Research · Knowledge an agent can cite ·

How the Earnings Lab is built

The Earnings Lab turns US quarterly earnings releases into an ontology-typed graph, answers questions over it as the signed-in user, and scores itself against the figures companies file in XBRL. This covers its parts on AWS and what it costs to keep and to run.

The Earnings Lab is an experiment in ontology-guided extraction. An agent reads each US company’s quarterly earnings release, extracts the facts it states into a graph typed by a hand-written earnings ontology, and a second agent answers questions over that graph. A harness then scores the answers.

The domain was chosen for one property most domains lack: a free ground truth. Every GAAP figure a company states in its press release is also filed in XBRL, and the SEC publishes those filings as the Financial Statement Data Sets. So extraction precision can be measured on every run, by comparing what the agent extracted with what the company filed, rather than by sampling and checking by hand.

The lab asks three questions. Can an ontology-guided agent turn earnings releases into facts that are correct, traceable and stable when a question is reworded? Does a typed graph earn its place over vector retrieval alone? And what does it cost to run over a whole quarter?

It is not production. It uses public SEC data only, runs in an isolated AWS account, and is torn down or paused when idle.

The parts

The lab runs left to right, from sources to the people asking questions.

Stage Component What it does
Sources SEC EDGAR 8-K filings carrying an earnings release, found through the form indexes
Sources Financial Statement Data Sets Quarterly XBRL facts for every 10-Q and 10-K: the ground truth
Sources Company tickers Ticker to SEC identifier map
Scan Ingest job (Fargate) Exhibit to text to passages, each with a stable id
Scan Data set and ticker jobs (Fargate) Quarterly zips to Parquet; ticker map to JSON lines
Scan Scheduler Daily ingest and weekly tickers, off until switched on
Model Earnings ontology OWL classes aligned to FIBO and US-GAAP, with SHACL shapes
Model Extraction agent (Fargate) Fills a tool schema generated from the ontology, with grounding checks
Model Graph loader Replaces each release’s named graph in Neptune
Model Governed metrics Each metric as an ordered list of XBRL concepts; no model involved
Stores S3 data lake Raw exhibits, text, passages, Parquet, staged graph files, run logs
Stores Glue and Athena Tables over the lake, with a per-query scan limit
Stores Amazon Neptune The ontology and the extracted facts, as RDF
Stores Bedrock Knowledge Base Passage vectors in S3 Vectors
Serve Graph tools (Lambda) Resolve terms to concepts, find a company, read facts, traverse, run SPARQL; read-only
Serve Data tools (Lambda) List and read governed metrics; retrieve passages by meaning
Interfaces AgentCore Gateway Exposes both sets of tools over MCP, accepting only user ID tokens
Consumers Q&A agent A Strands agent on AgentCore Runtime, acting for the caller, with conversation memory
Consumers Harness Asks families of paraphrased questions as named test users and scores them
Models Claude on Bedrock; Titan embeddings Extraction and answers; 1024-dimension passage vectors
Identity Cognito The single sign-in issuer; test users with different grants

How data flows

Onboarding the ground truth. The data set job fetches each quarter’s XBRL zip and writes Parquet by table and quarter. Athena reads it in place through partition projection, so nothing is loaded into a database.

Ingesting and extracting a quarter. The ingest job finds each quarter’s earnings filings, pulls the release exhibit, and splits it into passages. Each passage gets an id derived from the filing and a hash of its content, and is written as its own document for the Knowledge Base. The extraction agent then reads the passages, fills the ontology’s tool, and stages the facts that pass its checks as graph files. The loader replaces that release’s named graph in Neptune.

Answering a metric question. The harness signs in as a test user and asks the agent a question in several wordings. The agent calls the data tools through the Gateway with the user’s token. The tool resolves the metric to its ordered XBRL concepts and reads the value through Athena. The harness compares the answer with the governed value.

Answering a free-text question. The agent resolves the user’s terms to ontology concepts, reads matching facts and the passages they came from out of the graph, and retrieves similar passages by vector search. It answers from both and cites both.

Design decisions

The ontology is the contract. The tool the extraction model must answer through is generated from the ontology, so its allowed values are the ontology’s metric concepts, event types and guidance actions. Every extracted item cites the passage it was stated in. A number that does not appear as written, a quote that is not verbatim, or a SHACL violation goes back to the model as the tool result. After two repairs, failing items are dropped one at a time, so one bad fact never blocks a whole document.

One passage, one chunk. The Knowledge Base does no chunking of its own. A retrieved chunk is exactly one graph passage, with the passage id in its metadata, so vector results join the graph on that id without guessing at boundaries.

A named graph per release. Each release’s facts live in their own graph, so re-extracting a document replaces its facts atomically.

Metrics with no model in the loop. A metric answer comes from XBRL through a fixed list of concepts, including the derivation of a fourth quarter from year-to-date figures. That is what makes the harness’s correctness score mean something.

Identity on every hop. The user’s Cognito ID token travels through the Runtime and the Gateway to the tools, and memory is keyed by the caller. Machine credentials were rejected because they make answers unattributable.

Read-only tools. The agent’s tools can read the graph but have no write permission at the IAM level, so a question crafted to make the agent misuse a tool has nothing to write with.

Guardrails at the edges only. A Bedrock guardrail screens the user’s question and the final answer for harmful content and prompt attacks. It does not run on every model turn or on extraction: in an earlier setup it blocked the agent’s own narration and a filed press release, and those blocks looked like unrelated faults.

Store choices. S3 Vectors rather than OpenSearch Serverless, whose idle floor is USD 175 to 350 a month. Neptune Database on a small instance rather than Neptune Analytics, which supports openCypher only and the lab needs SPARQL. No NAT gateway, with Neptune private.

How it is tested

Check Passes when What it shows
Concept agreement Every paraphrase resolves to the same ontology concepts, and a near-miss question does not Meaning, not wording, drives retrieval
Evidence agreement The top five cited passages overlap across paraphrases Retrieval is stable under rewording
Correctness A metric answer is within 0.5% of the XBRL value The shared answer is also the right one
Extraction precision Extracted GAAP values match XBRL for the same company, metric and period The graph holds correct facts
Negative control An unanswerable question returns an explicit unknown The agent says when it does not know

Press releases also state non-GAAP figures with no XBRL counterpart. Those are checked for grounding in the text, but precision is reported on GAAP figures only.

What it costs

Prices below are AWS list prices in us-east-1, as researched in late September 2026.

Standing cost

Component Running, per month Paused or idle
Neptune (db.t4g.medium) about USD 68 storage only
S3, S3 Vectors, Glue, Athena a few dollars the same
Lambda, Fargate, AgentCore per use nothing
Cognito, secrets, parameters, logs about USD 2 the same
NAT gateway none, by design (saves about USD 33) none

Neptune is the only real standing cost. The lab’s control script pauses it when idle, though AWS restarts a stopped cluster after seven days.

Extraction cost

Extraction dominates. Measured with Claude Sonnet 4.5 on Bedrock, repairs included:

Run Documents Input tokens Output tokens Cost Per document
Earnings releases, 13 large companies 13 759,417 383,611 about USD 8.0 about USD 0.62
Matching earnings call transcripts 13 758,089 362,387 about USD 7.7 about USD 0.59

That is about 58,000 input and 29,000 output tokens a document: four times the input and ten times the output first assumed. Large-company releases are long and full of tables.

How cost scales

Scope a quarter Documents Sonnet on demand Haiku (about a third)
S&P 500 releases about 500 about USD 310 about USD 105
S&P 500 releases and calls about 960 about USD 580 about USD 195
Every US earnings release (unverified estimate) 5,000 to 7,000 USD 3,100 to 4,300 USD 1,000 to 1,450

Batch inference roughly halves any of these. For a rolling scope of about 450 large companies, one release and one call each, extraction on demand as configured measured about USD 0.94 a company, or about USD 430 a quarter. Batch inference with targeted repairs and trimmed output is estimated at about USD 70 a quarter, still to be confirmed by scoring. Once extraction is optimised, it is the smallest line, and the graph store (about USD 200 a quarter) decides what the scope costs.

Two things matter more than price once the lab grows:

Cost controls

A monthly budget on a project tag alerts at 50%, 80% and a forecast 100%. Athena caps each query’s scan. Schedules are off by default. Extraction takes a document limit, and the ingest manifest makes every run resumable rather than repeated. The chat sits behind access codes with daily limits and a global daily cap.

Limits

Both test users currently reach the same tools, so a negative entitlement check cannot yet run on this path. The precision measure depends on XBRL, so the shape does not carry over to a domain without a structured ground truth.