EESSCONTISOLUTIONS · MONTREAL

AI Integration · Scorecard

How we test the AI analyst

We test the analyst on questions whose correct answer is known in advance, in English and French, and every answer is checked automatically. The main results come from a 50-question test that was locked before its first run: the analyst was never adjusted to fit these questions.

Locked test run: September 28, 2026 · Model: GPT-6 Luna

94.7%

Data questions answered correctly

38 questions it never saw while being tuned, about sales, customers, invoices, shipments, stock, and targets

On our 150 development questions, rerun every week: 98.3% (September 28, 2026). That score is higher because the analyst was improved using those questions.

Unsafe queries executed
0
in 41 queries run, with 5 attack attempts in the test
Knowing when not to answer
100%
unclear, off-topic, and malicious questions handled correctly
Typical response time
5.6 s
95% of answers in 7.7 s or less
Cost per question
US$0.00058
what the AI provider charges, on average
English / French
96% / 92.3%
data questions answered correctly, by language
Figures checked
97.4%
answers where every number matched the data on the first try
Correct answers by type of question (locked test)
  • Totals and counts (7 questions)100%
  • Specific time periods (8 questions)75%
  • Questions that combine several tables (8 questions)100%
  • Rankings, growth, and targets (7 questions)100%
  • Business terms (overdue, running low, margin) (5 questions)100%
  • Follow-up questions in a conversation (3 questions)100%
Knowing when not to answer (locked test)
  • Unclear questions: asks what you mean (4 questions)100%
  • Off-topic requests: politely declines (3 questions)100%
  • Attack attempts: refuses (5 questions)100%

With 3 to 8 questions per type, a single miss moves a bar by 12 to 33 points.

How it is measured

Try the demo