69 pages of bureaucracy, with two different ceilings that look like the same thing.
Run on 8 October 2026
- correct answers
- 25 / 25
- points on the protocol's rubric
- 50 / 50
- fabricated answers
- 0
Retrieval quality on Romanian administrative prose was this product’s largest untested assumption. It is a failure mode that does not announce itself: a system that is excellent in English degrades quietly on another language — slightly worse ranking, slightly thinner answers — and nobody in this market measures it. So it was measured.
The protocol was written before the run
Twenty-five questions across five categories — plain factual lookup, traps, lists and definitions, synthesis, and questions about the document as a whole — each with its correct answer recorded in advance and verified by hand against the guide. Only then did the system see the document.
That ordering is the entire difference between a measurement and a demo. A demo shows you the questions that went well.
The guide sat alone, in a test company created for the run. In the August run it sat beside another company’s internal documents, and “Is this document final?” had no clear referent — the only point lost then.
| Score | Meaning |
|---|---|
| 2 | Correct, and the citation points at the right section |
| 1 | Correct, but the citation is missing or wrong |
| 0 | Wrong, incomplete, or “I don't know” where an answer exists |
| −2 | Fabricated — a fact that is not in the document |
The pass mark was fixed in advance too: a mean of at least 1.5, and zero −2 scores on the trap questions. One invented figure in a document the reader knows by heart costs you the account, so that column has no tolerance.
And the rubric is applied strictly, to the whole answer: a correct answer that adds, on the side, a fact the guide does not hold scores −2, not 2. To the reader, that fact looks exactly like one from the document.
The four questions it turns on
The answers below are the literal output, read back out of the database. They are quoted in Romanian because that is what makes them evidence; the English gloss follows each one.
The discriminator
What is the maximum non-reimbursable funding that can be requested?
Valoarea maximă a finanțării nerambursabile solicitate este de 200.000 euro. În cadrul acestui apel, finanțarea nerambursabilă solicitată poate fi de maximum 200.000 euro, cu respectarea condițiilor aplicabile ajutoarelor de minimis. […] Valoarea eligibilă a proiectului: minimum 20.000 euro și maximum 300.000 euro.
“At most €200,000, subject to de minimis rules. […] Eligible project value: at least €20,000 and at most €300,000.” The €300,000 is a different ceiling that looks identical at a glance, and keyword matching returns it. It returned the right one and set the other apart, labelled.
Hallucination trap
What is the submission deadline?
Ghidul nu indică o dată concretă pentru termenul limită. Secțiunea 2.5 „Perioada de depunere” are câmpurile necompletate: data deschiderii apelului, data și ora de începere a depunerii și data și ora de închidere a depunerii. Pentru data exactă trebuie consultat anunțul oficial de lansare a apelului sau versiunea finală a ghidului.
“The guide gives no deadline. Section 2.5, ‘Submission period’, leaves its fields blank: the opening date, and the start and closing date and time. For the exact date, see the call’s official announcement or the final guide.” The value is not there. It named the section and the empty fields instead of producing a plausible date.
The second trap
Which exchange rate applies for conversion into lei?
Ghidul folosește cursul InforEuro, dar luna de referință diferă după scopul conversiei. […] Textul ghidului din rezultatele căutării are câmpuri necompletate pentru luna și valoarea cursului (1 euro = … lei). Nu pot indica un curs numeric concret din aceste fragmente.
“The guide uses the InforEuro rate, but the reference month depends on the purpose. […] The month and the rate are left blank (1 euro = … lei), so I cannot give a figure.” The draft leaves the rate empty. It gave the rule, tabled the five cases where the guide fixes a different month, and gave no number.
Reasoning
An SME wants a €250,000 project. How much can it receive, and how much must it put in?
Pentru un proiect de 250.000 euro, presupunând că toată suma este eligibilă, IMM-ul poate primi cel mult 200.000 euro nerambursabili. Trebuie să pună cel puțin 50.000 euro. […] Plafonul de 200.000 euro se aplică înaintea procentului de 90%. De aceea contribuția de 10% (25.000 euro) nu ajunge, iar diferența până la 250.000 euro cade tot în sarcina solicitantului.
“At most €200,000; it must put in at least €50,000. […] The €200,000 ceiling applies before the 90% rate, so a 10% contribution (€25,000) is not enough.” This answer is written nowhere in the guide. It requires two rules from different sections — the 90% rate and the de minimis ceiling — plus noticing that the second overrides the first.
How it got to 25 out of 25
The first run on 8 October 2026, on the product as it was that morning, scored 41 out of 50. All 25 answers were right in substance, but two brought in, from the knowledge base, examples the guide does not contain: “Google Workspace, Microsoft 365” for SaaS, sensors measuring temperature and pressure for IoT. The compiler that turns documents into knowledge pages had written them — from what it knew about the world, not from the guide.
To the reader that looks exactly like information from the document. The compiler was fixed — every statement and figure on a page must come from its source — and the run repeated with the same questions and the same rubric. The figures at the top are that run.
One note, without a deduction: on the €250,000 project, the answer places a contribution of exactly 20% in the bracket worth 4 points, although the guide’s grid puts 20% in two brackets at once. Not an invented fact, but a certainty the guide does not give.
The bigger test: a whole company
A long guide is a reading test. The product promises more: one AI across the whole company — policies, contracts, invoices, email, Slack, Notion, GitHub — where each person sees only what they may. So a complete test company was built, and 55 questions, written before the run, were asked as four different people: the owner, an admin, and two employees on different teams.
- leaks across access levels, on 11 permission questions
- 0
- actions run without a person's approval
- 0
- values invented on the 8 trap questions
- 0
On the rubric, 104 out of 110. One answer said something wrong: it repeated a knowledge page that had linked one person’s role to another person with a similar name. That was fixed too.
A second run, on the shipped version, stopped after 50 questions and scored the same on them, with two different mistakes: a weekday worked out wrong in an aside nobody asked for, and a refusal to show the owner a salary from their own document, because the document said assistants must refuse. A single run’s score moves by a few points. The three figures in the box above held in every run.
Method, and what the test does not show
The document is the applicant’s guide for Action 2.2 of the Regiunea Centru 2021–2027 programme, published by the regional development agency: 69 pages, marked DRAFT 1, public and downloadable. The run was made on 8 October 2026 through the product’s API, across five separate conversations, one per section of the protocol — on the code that then went to production, in a test database, not on any client’s data. The shipped version adds three adjustments made after the guide’s run (reading a whole file, which source governs when two disagree, counting); the guide was not re-run on it. Every answer, citation and grounding score is kept and can be re-read.
What it does not show: that the figures in this guide are current. The tested document is a draft of an earlier call, so the eligibility values appearing in the answers above are historical. The test says something about the system, not about any live call.
Nor does it show that the system never errs. It shows that on a real Romanian document, faced with two questions built specifically to make it invent, it did not.
Ask any AI vendor the same question.
Not which model they use — whether they have a test protocol in the language their users actually write in, written before the run, and what it scored.
hello@temiro.ai