Research · Knowledge an agent can cite ·
How the Earnings Lab is built
The Earnings Lab turns US quarterly earnings releases into an ontology-typed graph, answers questions over it as the signed-in user, and scores itself against the figures companies file in XBRL. This covers its parts on AWS and what it costs to keep and to run.
The Earnings Lab is an experiment in ontology-guided extraction. An agent reads each US company’s quarterly earnings release, extracts the facts it states into a graph typed by a hand-written earnings ontology, and a second agent answers questions over that graph. A harness then scores the answers.
The domain was chosen for one property most domains lack: a free ground truth. Every GAAP figure a company states in its press release is also filed in XBRL, and the SEC publishes those filings as the Financial Statement Data Sets. So extraction precision can be measured on every run, by comparing what the agent extracted with what the company filed, rather than by sampling and checking by hand.
The lab asks three questions. Can an ontology-guided agent turn earnings releases into facts that are correct, traceable and stable when a question is reworded? Does a typed graph earn its place over vector retrieval alone? And what does it cost to run over a whole quarter?
It is not production. It uses public SEC data only, runs in an isolated AWS account, and is torn down or paused when idle.
The parts
The lab runs left to right, from sources to the people asking questions.
| Stage | Component | What it does |
|---|---|---|
| Sources | SEC EDGAR | 8-K filings carrying an earnings release, found through the form indexes |
| Sources | Financial Statement Data Sets | Quarterly XBRL facts for every 10-Q and 10-K: the ground truth |
| Sources | Company tickers | Ticker to SEC identifier map |
| Scan | Ingest job (Fargate) | Exhibit to text to passages, each with a stable id |
| Scan | Data set and ticker jobs (Fargate) | Quarterly zips to Parquet; ticker map to JSON lines |
| Scan | Scheduler | Daily ingest and weekly tickers, off until switched on |
| Model | Earnings ontology | OWL classes aligned to FIBO and US-GAAP, with SHACL shapes |
| Model | Extraction agent (Fargate) | Fills a tool schema generated from the ontology, with grounding checks |
| Model | Graph loader | Replaces each release’s named graph in Neptune |
| Model | Governed metrics | Each metric as an ordered list of XBRL concepts; no model involved |
| Stores | S3 data lake | Raw exhibits, text, passages, Parquet, staged graph files, run logs |
| Stores | Glue and Athena | Tables over the lake, with a per-query scan limit |
| Stores | Amazon Neptune | The ontology and the extracted facts, as RDF |
| Stores | Bedrock Knowledge Base | Passage vectors in S3 Vectors |
| Serve | Graph tools (Lambda) | Resolve terms to concepts, find a company, read facts, traverse, run SPARQL; read-only |
| Serve | Data tools (Lambda) | List and read governed metrics; retrieve passages by meaning |
| Interfaces | AgentCore Gateway | Exposes both sets of tools over MCP, accepting only user ID tokens |
| Consumers | Q&A agent | A Strands agent on AgentCore Runtime, acting for the caller, with conversation memory |
| Consumers | Harness | Asks families of paraphrased questions as named test users and scores them |
| Models | Claude on Bedrock; Titan embeddings | Extraction and answers; 1024-dimension passage vectors |
| Identity | Cognito | The single sign-in issuer; test users with different grants |
How data flows
Onboarding the ground truth. The data set job fetches each quarter’s XBRL zip and writes Parquet by table and quarter. Athena reads it in place through partition projection, so nothing is loaded into a database.
Ingesting and extracting a quarter. The ingest job finds each quarter’s earnings filings, pulls the release exhibit, and splits it into passages. Each passage gets an id derived from the filing and a hash of its content, and is written as its own document for the Knowledge Base. The extraction agent then reads the passages, fills the ontology’s tool, and stages the facts that pass its checks as graph files. The loader replaces that release’s named graph in Neptune.
Answering a metric question. The harness signs in as a test user and asks the agent a question in several wordings. The agent calls the data tools through the Gateway with the user’s token. The tool resolves the metric to its ordered XBRL concepts and reads the value through Athena. The harness compares the answer with the governed value.
Answering a free-text question. The agent resolves the user’s terms to ontology concepts, reads matching facts and the passages they came from out of the graph, and retrieves similar passages by vector search. It answers from both and cites both.
Design decisions
The ontology is the contract. The tool the extraction model must answer through is generated from the ontology, so its allowed values are the ontology’s metric concepts, event types and guidance actions. Every extracted item cites the passage it was stated in. A number that does not appear as written, a quote that is not verbatim, or a SHACL violation goes back to the model as the tool result. After two repairs, failing items are dropped one at a time, so one bad fact never blocks a whole document.
One passage, one chunk. The Knowledge Base does no chunking of its own. A retrieved chunk is exactly one graph passage, with the passage id in its metadata, so vector results join the graph on that id without guessing at boundaries.
A named graph per release. Each release’s facts live in their own graph, so re-extracting a document replaces its facts atomically.
Metrics with no model in the loop. A metric answer comes from XBRL through a fixed list of concepts, including the derivation of a fourth quarter from year-to-date figures. That is what makes the harness’s correctness score mean something.
Identity on every hop. The user’s Cognito ID token travels through the Runtime and the Gateway to the tools, and memory is keyed by the caller. Machine credentials were rejected because they make answers unattributable.
Read-only tools. The agent’s tools can read the graph but have no write permission at the IAM level, so a question crafted to make the agent misuse a tool has nothing to write with.
Guardrails at the edges only. A Bedrock guardrail screens the user’s question and the final answer for harmful content and prompt attacks. It does not run on every model turn or on extraction: in an earlier setup it blocked the agent’s own narration and a filed press release, and those blocks looked like unrelated faults.
Store choices. S3 Vectors rather than OpenSearch Serverless, whose idle floor is USD 175 to 350 a month. Neptune Database on a small instance rather than Neptune Analytics, which supports openCypher only and the lab needs SPARQL. No NAT gateway, with Neptune private.
How it is tested
| Check | Passes when | What it shows |
|---|---|---|
| Concept agreement | Every paraphrase resolves to the same ontology concepts, and a near-miss question does not | Meaning, not wording, drives retrieval |
| Evidence agreement | The top five cited passages overlap across paraphrases | Retrieval is stable under rewording |
| Correctness | A metric answer is within 0.5% of the XBRL value | The shared answer is also the right one |
| Extraction precision | Extracted GAAP values match XBRL for the same company, metric and period | The graph holds correct facts |
| Negative control | An unanswerable question returns an explicit unknown | The agent says when it does not know |
Press releases also state non-GAAP figures with no XBRL counterpart. Those are checked for grounding in the text, but precision is reported on GAAP figures only.
What it costs
Prices below are AWS list prices in us-east-1, as researched in late September 2026.
Standing cost
| Component | Running, per month | Paused or idle |
|---|---|---|
| Neptune (db.t4g.medium) | about USD 68 | storage only |
| S3, S3 Vectors, Glue, Athena | a few dollars | the same |
| Lambda, Fargate, AgentCore | per use | nothing |
| Cognito, secrets, parameters, logs | about USD 2 | the same |
| NAT gateway | none, by design (saves about USD 33) | none |
Neptune is the only real standing cost. The lab’s control script pauses it when idle, though AWS restarts a stopped cluster after seven days.
Extraction cost
Extraction dominates. Measured with Claude Sonnet 4.5 on Bedrock, repairs included:
| Run | Documents | Input tokens | Output tokens | Cost | Per document |
|---|---|---|---|---|---|
| Earnings releases, 13 large companies | 13 | 759,417 | 383,611 | about USD 8.0 | about USD 0.62 |
| Matching earnings call transcripts | 13 | 758,089 | 362,387 | about USD 7.7 | about USD 0.59 |
That is about 58,000 input and 29,000 output tokens a document: four times the input and ten times the output first assumed. Large-company releases are long and full of tables.
How cost scales
| Scope a quarter | Documents | Sonnet on demand | Haiku (about a third) |
|---|---|---|---|
| S&P 500 releases | about 500 | about USD 310 | about USD 105 |
| S&P 500 releases and calls | about 960 | about USD 580 | about USD 195 |
| Every US earnings release (unverified estimate) | 5,000 to 7,000 | USD 3,100 to 4,300 | USD 1,000 to 1,450 |
Batch inference roughly halves any of these. For a rolling scope of about 450 large companies, one release and one call each, extraction on demand as configured measured about USD 0.94 a company, or about USD 430 a quarter. Batch inference with targeted repairs and trimmed output is estimated at about USD 70 a quarter, still to be confirmed by scoring. Once extraction is optimised, it is the smallest line, and the graph store (about USD 200 a quarter) decides what the scope costs.
Two things matter more than price once the lab grows:
- Re-extraction. Every ontology change that alters what is extracted means running extraction again. Three revisions over a 12-quarter backfill of the S&P 500 would cost about USD 22,000 on Sonnet, so the ontology is settled on a small paired set first.
- Account limits. A new AWS account has a daily Bedrock token allowance and no batch inference until AWS Support verifies it. One quarter of S&P 500 releases and calls is about 85 million tokens, roughly two weeks at the allowance seen on day one.
Cost controls
A monthly budget on a project tag alerts at 50%, 80% and a forecast 100%. Athena caps each query’s scan. Schedules are off by default. Extraction takes a document limit, and the ingest manifest makes every run resumable rather than repeated. The chat sits behind access codes with daily limits and a global daily cap.
Limits
Both test users currently reach the same tools, so a negative entitlement check cannot yet run on this path. The precision measure depends on XBRL, so the shape does not carry over to a domain without a structured ground truth.