Research · Knowledge an agent can cite ·
How the agents are tested
The Earnings Lab has two agents, one that reads earnings releases into a graph and one that answers questions over it. This covers how each is checked, what happens to a failure, and what is not tested yet.
The Earnings Lab runs two agents. The extraction agent reads each company’s quarterly earnings release and writes the facts it states into a graph typed by the earnings ontology. The chat agent answers questions over that graph and cites what it used. Both can be wrong in ways that read well, so each has its own reference to be checked against.
| What is tested | Against what | When |
|---|---|---|
| Every fact is grounded in its source | The cited passage of the release | Every extraction, before anything is stored |
| Every fact fits the ontology | SHACL shapes generated beside the ontology | Every extraction, before anything is stored |
| Extracted figures are correct | The XBRL data companies file with the SEC | Every scored extraction run |
| Chat answers are correct and cited | XBRL, the source text, and hand-written answer keys | Eval runs, after a change |
Checks on every extraction
The extraction agent cannot answer in free text. It must answer through a tool whose schema is generated from the ontology, so it can only emit what the ontology defines, and every item must cite its passages. The result then passes three deterministic gates.
- Shape. A missing unit or a period without a fiscal year is caught first.
- Grounding. Every cited passage must exist. Every quote must appear verbatim in its cited passages. Every reported figure and every guidance bound must appear in its cited passages as a whole number as written, allowing for rounding in either direction but never matching a fragment: 12 does not match “12,400”.
- Ontology conformance. The facts are built into RDF and validated with SHACL against the ontology. A reported value needs exactly one company, metric, period, value, unit and basis (GAAP or non-GAAP). Guidance cannot have a low above its high. Every fact must name its passage.
What a failure looks like
A failed check produces a message the agent can act on, such as “reported_values[3] (Revenue = 4210) does not appear in its cited passages; cite the passage that states it, or correct the number”. The messages go back to the agent for a repair turn, twice at most. Whatever still fails after that is dropped item by item, so one bad figure does not cost the whole release.
Nothing is dropped silently. Each stored extraction keeps the errors of every rejected turn and each dropped item with the errors that dropped it. A release that fails outright is kept beside its error, retried once, then tried on a fallback model. A data quality report counts, per run, the documents that were clean, repaired, partly dropped or failed.
Scoring extraction against XBRL
Earnings releases have a property most documents lack: the GAAP figures they state are also filed in XBRL, and the SEC publishes those filings as structured data sets. The harness compares every consolidated GAAP figure in a run with the filed value for the same company, metric and quarter.
- Precision counts the extracted figures that match the filed one, within 0.5% or within the rounding the release stated it at (“$2.2 million” matches anything that rounds to it).
- Recall counts the filed figures for the release’s own quarter that its text states, and how many of them extraction got right.
- Figures with no filed counterpart are reported as unscorable, not as correct.
The same scoring compares models. On 11 releases, the model now used for releases was as precise against XBRL as two Claude models and found more of the stated GAAP figures, but failed outright on 2 of them, hence the fallback. Samples this size show a clearly worse model, not a difference of a few points.
Testing the chat agent
Each test question goes through the published chat with an access token and the answer cache off, as a user would ask it.
The main set is 41 competency questions: questions the ontology must be able to answer, written before looking at extraction results, so the ontology is judged against what people want to ask rather than tuned to what the model extracted. They cover results, outlook, explanations, analysts’ questions, people, consistency between release and call, events, change over time, refusals and privacy, and each names its source of truth.
| Grader | How it decides | Questions |
|---|---|---|
| Numeric | The figure in the answer against XBRL | 4 |
| Rubric | A model judge reads the release and call text | 25 |
| Set | The judge checks a list against the sources | 8 |
| Refusal | The answer must say the lab does not hold it | 2 |
| Canary | A guest answer must not contain any private-only fact | 2 |
The judge grades against the source documents, never its own knowledge of the company. It judges separately whether the answer is right and whether the graph held the facts: an answer can pass from the release text while the graph lacks the fact, and that gap drives ontology work. Every failure is attributed to one cause, and only one argues for changing the ontology.
| Attribution | Meaning |
|---|---|
| Ontology | No class, property or value exists to hold the answer |
| Extraction | The term exists and the source states it, but the graph lacks the fact |
| Retrieval | The fact is in the graph but the agent’s queries missed it |
| Answer | The fact was found but misstated, or the answer is uncited |
When the cause is the ontology, the judge proposes the missing terms. A person decides whether to add them.
A golden set of paraphrase families asks one meaning in several wordings (“revenue”, “net sales”, “top line”) as test users with different permissions. It scores whether every wording resolves to the same concepts, cites overlapping evidence and gives the right figure, read from XBRL at run time so it cannot drift. A held-out set of 59 questions checks that the agent’s router picks a suitable playbook for each kind of question.
What is not tested yet
- Facts from earnings call transcripts are counted, not scored. Nothing compares them with a reference yet, so a model comparison on calls shows what each model found, not whether it was right.
- XBRL covers consolidated GAAP figures only. Non-GAAP figures, segment figures, guidance, drivers, risks and quotes are checked for grounding and shape, but their correctness rests on the model judge.
- Grounding proves a number appears in the cited passage. It does not prove the number was attached to the right metric or period. XBRL catches that for GAAP figures. For the rest, only the model judge can, and only for the questions it is asked.
- The judge is a model and can be wrong. It reads a bounded amount of each document, and a rubric grade is a model’s reading of another model’s answer.
- Questions about change over time need a second quarter of data and are expected to fail until it is loaded. Some golden families still lack hand-checked expected passages.
- The competency questions and the router’s held-out set are drafts, written alongside the system. Questions from real use will replace or extend them.