Research · Knowledge an agent can cite ·

How the agents are tested

The Earnings Lab has two agents, one that reads earnings releases into a graph and one that answers questions over it. This covers how each is checked, what happens to a failure, and what is not tested yet.

The Earnings Lab runs two agents. The extraction agent reads each company’s quarterly earnings release and writes the facts it states into a graph typed by the earnings ontology. The chat agent answers questions over that graph and cites what it used. Both can be wrong in ways that read well, so each has its own reference to be checked against.

What is tested Against what When
Every fact is grounded in its source The cited passage of the release Every extraction, before anything is stored
Every fact fits the ontology SHACL shapes generated beside the ontology Every extraction, before anything is stored
Extracted figures are correct The XBRL data companies file with the SEC Every scored extraction run
Chat answers are correct and cited XBRL, the source text, and hand-written answer keys Eval runs, after a change

Checks on every extraction

The extraction agent cannot answer in free text. It must answer through a tool whose schema is generated from the ontology, so it can only emit what the ontology defines, and every item must cite its passages. The result then passes three deterministic gates.

  1. Shape. A missing unit or a period without a fiscal year is caught first.
  2. Grounding. Every cited passage must exist. Every quote must appear verbatim in its cited passages. Every reported figure and every guidance bound must appear in its cited passages as a whole number as written, allowing for rounding in either direction but never matching a fragment: 12 does not match “12,400”.
  3. Ontology conformance. The facts are built into RDF and validated with SHACL against the ontology. A reported value needs exactly one company, metric, period, value, unit and basis (GAAP or non-GAAP). Guidance cannot have a low above its high. Every fact must name its passage.

What a failure looks like

A failed check produces a message the agent can act on, such as “reported_values[3] (Revenue = 4210) does not appear in its cited passages; cite the passage that states it, or correct the number”. The messages go back to the agent for a repair turn, twice at most. Whatever still fails after that is dropped item by item, so one bad figure does not cost the whole release.

Nothing is dropped silently. Each stored extraction keeps the errors of every rejected turn and each dropped item with the errors that dropped it. A release that fails outright is kept beside its error, retried once, then tried on a fallback model. A data quality report counts, per run, the documents that were clean, repaired, partly dropped or failed.

Scoring extraction against XBRL

Earnings releases have a property most documents lack: the GAAP figures they state are also filed in XBRL, and the SEC publishes those filings as structured data sets. The harness compares every consolidated GAAP figure in a run with the filed value for the same company, metric and quarter.

The same scoring compares models. On 11 releases, the model now used for releases was as precise against XBRL as two Claude models and found more of the stated GAAP figures, but failed outright on 2 of them, hence the fallback. Samples this size show a clearly worse model, not a difference of a few points.

Testing the chat agent

Each test question goes through the published chat with an access token and the answer cache off, as a user would ask it.

The main set is 41 competency questions: questions the ontology must be able to answer, written before looking at extraction results, so the ontology is judged against what people want to ask rather than tuned to what the model extracted. They cover results, outlook, explanations, analysts’ questions, people, consistency between release and call, events, change over time, refusals and privacy, and each names its source of truth.

Grader How it decides Questions
Numeric The figure in the answer against XBRL 4
Rubric A model judge reads the release and call text 25
Set The judge checks a list against the sources 8
Refusal The answer must say the lab does not hold it 2
Canary A guest answer must not contain any private-only fact 2

The judge grades against the source documents, never its own knowledge of the company. It judges separately whether the answer is right and whether the graph held the facts: an answer can pass from the release text while the graph lacks the fact, and that gap drives ontology work. Every failure is attributed to one cause, and only one argues for changing the ontology.

Attribution Meaning
Ontology No class, property or value exists to hold the answer
Extraction The term exists and the source states it, but the graph lacks the fact
Retrieval The fact is in the graph but the agent’s queries missed it
Answer The fact was found but misstated, or the answer is uncited

When the cause is the ontology, the judge proposes the missing terms. A person decides whether to add them.

A golden set of paraphrase families asks one meaning in several wordings (“revenue”, “net sales”, “top line”) as test users with different permissions. It scores whether every wording resolves to the same concepts, cites overlapping evidence and gives the right figure, read from XBRL at run time so it cannot drift. A held-out set of 59 questions checks that the agent’s router picks a suitable playbook for each kind of question.

What is not tested yet