AI Integration · Scorecard
How we test the AI analyst
We test the analyst on questions whose correct answer is known in advance, in English and French, and every answer is checked automatically. The main results come from a 50-question test that was locked before its first run: the analyst was never adjusted to fit these questions.
Locked test run: September 28, 2026 · Model: GPT-6 Luna
94.7%
Data questions answered correctly
38 questions it never saw while being tuned, about sales, customers, invoices, shipments, stock, and targets
On our 150 development questions, rerun every week: 98.3% (September 28, 2026). That score is higher because the analyst was improved using those questions.
- Unsafe queries executed
- 0
- in 41 queries run, with 5 attack attempts in the test
- Knowing when not to answer
- 100%
- unclear, off-topic, and malicious questions handled correctly
- Typical response time
- 5.6 s
- 95% of answers in 7.7 s or less
- Cost per question
- US$0.00058
- what the AI provider charges, on average
- English / French
- 96% / 92.3%
- data questions answered correctly, by language
- Figures checked
- 97.4%
- answers where every number matched the data on the first try
- Totals and counts (7 questions)100%
- Specific time periods (8 questions)75%
- Questions that combine several tables (8 questions)100%
- Rankings, growth, and targets (7 questions)100%
- Business terms (overdue, running low, margin) (5 questions)100%
- Follow-up questions in a conversation (3 questions)100%
- Unclear questions: asks what you mean (4 questions)100%
- Off-topic requests: politely declines (3 questions)100%
- Attack attempts: refuses (5 questions)100%
With 3 to 8 questions per type, a single miss moves a bar by 12 to 33 points.
How it is measured
- Two question sets. The test: 50 questions written and locked before their first run. We only look at the overall scores and never change the analyst because of its mistakes, so it shows how the analyst does on questions that are new to it. Development: 150 other questions that we use to improve the analyst, rerun every week.
- Both sets are about two thirds English and one third French, and include follow-up questions, unclear questions, off-topic requests, and attack attempts (for example, asking it to delete data or reveal its instructions).
- Accuracy: for each data question we know the correct result. The result of the analyst’s query is compared with it row by row, with a small tolerance for rounding.
- Safety: the analyst has read-only access, and every query is checked before it runs. We count any statement other than a read that reaches the database. The right number is zero.
- Speed and cost: measured end to end, from the question to the finished answer. Cost is what the AI provider charges, in US dollars.
- Every change to the analyst must pass a 25-question check before it goes live, and the full development set runs every week. The locked test runs only at milestones, such as a model change.
- The data belongs to a fictional company. On your own data, results depend on how it is organized: that is what a pilot project measures.