Reproducible test

Benchmarked against a real 69-page funding guide.

Run in production · 25 August 2026

correct answers
24 / 25
points on the protocol's rubric
48 / 50
fabricated answers
0

Retrieval quality on Romanian administrative prose was this product’s largest untested assumption. It is a failure mode that does not announce itself: a system that is excellent in English degrades quietly on another language — slightly worse ranking, slightly thinner answers — and nobody in this market measures it. So it was measured.

The protocol was written before the run

Twenty-five questions across five categories — plain factual lookup, traps, lists and definitions, synthesis, and questions about the document as a whole — each with its correct answer recorded in advance and verified by hand against the guide. Only then did the system see the document.

That ordering is the entire difference between a measurement and a demo. A demo shows you the questions that went well.

The guide was also not alone in the index. Sitting beside it were the company’s own twenty internal documents — contracts, policies, rate cards — so every question competed against a second corpus. Two of the results below are explained by exactly that.

ScoreMeaning
2Correct, and the citation points at the right section
1Correct, but the citation is missing or wrong
0Wrong, incomplete, or “I don't know” where an answer exists
−2Fabricated — a fact that is not in the document

The pass mark was fixed in advance too: a mean of at least 1.5, and zero −2 scores on the trap questions. One invented figure in a document the reader knows by heart costs you the account, so that column has no tolerance.

The four questions it turns on

The answers below are the literal output, read back out of the database. They are quoted in Romanian because that is what makes them evidence; the English gloss follows each one.

The discriminator

What is the maximum non-reimbursable funding that can be requested?

Valoarea finanțării nerambursabile solicitate poate fi de maximum 200.000 euro, cu respectarea condițiilor aplicabile ajutoarelor de minimis.

“At most €200,000, subject to de minimis rules.” The document also contains €300,000 — the maximum eligible project value, a different ceiling that looks identical at a glance. Keyword matching returns the wrong one. It returned the right one and then explained the difference between the two unprompted.

Hallucination trap

What is the submission deadline?

Data și ora de închidere a depunerii de proiecte: ............................ — aceste date sunt lăsate ca spații libere/placeholder, deci nu pot extrage din document o dată-limită concretă.

“Closing date and time: …………… — these are left as blanks, so I cannot extract a concrete deadline from the document.” The value simply is not there. It quoted the literal ellipses out of the guide and pointed at the agency's website for the final version, rather than producing a plausible date.

Two sources that disagree

Which exchange rate applies for conversion into lei?

Dacă te referi la facturarea clienților studioului, e cursul BNR. Dacă te referi la un proiect cu finanțare nerambursabilă, e cursul InforEuro din luna lansării apelului.

“For invoicing the studio's clients it is the national bank rate; for a grant-funded project it is the InforEuro rate for the month the call opened.” The knowledge base held both the guide and the company's own invoicing policy — two correct answers to one question. It neither blended them nor silently picked one: two labelled sections, each with its own citations, and a question back about which context was meant.

Reasoning

An SME wants a €250,000 project. How much can it receive, and how much must it put in?

IMM-ul nu poate primi 225.000 euro (adică 90% din 250.000), ci maxim 200.000 euro. Contribuție proprie: 50.000 euro.

“Not €225,000 (90% of 250,000) but at most €200,000; own contribution €50,000.” This answer is written nowhere in the guide. It requires two rules from different sections — the 90% co-financing rate and the de minimis ceiling — plus noticing that the second overrides the first.

The question it got wrong

One in twenty-five, reported here as prominently as the wins. Asked “Is this document final?” — the guide is marked DRAFT 1 — it replied “I don’t have enough context; which document do you mean?”, with no citation. On the rubric that is a zero.

The cause is the test condition itself. The knowledge base held not only the guide but the company’s own twenty internal documents, and “this document”, in a conversation that opens with that question, has no unique referent. It asked rather than guessed.

There is a good argument that this is the correct behaviour, and it probably is. But the protocol was written before the run and the question does have an answer in the document, so it scores zero — you do not rewrite the rubric once you have seen the result. That is the difference between a measurement and a case for the defence.

Method, and what the test does not show

The document is the applicant’s guide for Action 2.2 of the Regiunea Centru 2021–2027 programme, published by the regional development agency: 69 pages, marked DRAFT 1, public and downloadable. The run was made in production on 25 August 2026 across five separate conversations, one per section of the protocol. Every answer, citation and grounding score is still in the database and can be re-read.

What it does not show: that the figures in this guide are current. The tested document is a draft of an earlier call, so the eligibility values appearing in the answers above are historical. The test says something about the system, not about any live call.

Nor does it show that the system never errs. It shows that on a real Romanian document, faced with two questions built specifically to make it invent, it did not.

30-minute call

Ask any AI vendor the same question.

Not which model they use — whether they have a test protocol in the language their users actually write in, written before the run, and what it scored.

hello@temiro.ai