Guide · Start here

An architect's view

In a nutshell

For a technical audience: enterprise and solution architects. The lab is a working reference for a governed semantic layer for AI: an ontology as the single contract, verified extraction with provenance, agents acting as the user through MCP, evaluation as a component, and cost, licensing and abuse controls built in.

For a technical audience. This article is written for enterprise and solution architects, and assumes familiarity with knowledge graphs, retrieval-augmented generation and cloud platforms. For what the lab does in plain terms, read about the earnings lab.

The problem it addresses

Most organisations hold their knowledge in documents, and most generative AI pilots over those documents stall at the same point: the answers sound right, but nobody can say where they came from, whether they are complete, or what they cost. The lab treats that as an architecture problem, not a prompting problem. It builds a governed semantic layer between the documents and the AI, and tests it end to end on SEC earnings filings, where XBRL data provides a free ground truth to score against.

Each concept below is implemented and running, not drawn on a slide.

1. The ontology is the contract

One OWL ontology (the ontology), aligned to FIBO and the US-GAAP taxonomy, drives every layer:

Generated from the ontology Purpose
The extraction tool's JSON schema The model can only emit what the ontology defines
SHACL shapes A validation gate before anything reaches the graph
Synonyms and XBRL alignments Resolve a user's wording to governed concepts
The agent's term resolution One vocabulary for extraction and for questions
This page Every node typed and explained by its class

Changing the model of the domain is one change, versioned and reviewed, rather than a hunt through prompts and schemas that drift apart.

2. Extraction you can verify

Extraction runs force structured output through a tool, then verify it deterministically: every number must appear in its cited passage, quotes must be verbatim, and each fact must pass SHACL. Failures go back to the model for at most two repair turns, and whatever still fails is dropped item by item rather than failing the document. Every fact records its passage and its run (model id and prompt version) using W3C PROV, and each release loads into its own named graph, so reloads are idempotent and runs can be compared.

3. Graph and retrieval share one identifier

A passage has one stable id, used both as the chunk id in the vector index and as a node in the graph. That single join is what makes hybrid retrieval coherent: a semantic search hit leads straight to the structured facts extracted from it, and a graph answer leads back to the text that supports it.

4. Governed metrics beside extracted facts

Headline metrics are defined once, in a governed metric catalogue, and resolved against the SEC's Financial Statement Data Sets queried through Athena. The agent can tell a figure from a system of record apart from one extracted from prose, and the harness scores extraction against the former.

5. Agents that act as the user, through MCP

The question-answering agent runs on Amazon Bedrock AgentCore and reaches its tools through an AgentCore Gateway over the Model Context Protocol, acting as the signed-in user rather than as a service account. Tool design proved as important as model choice. Narrow, fine-grained tools are kept for the evaluation harness; the public agent gets coarse analyst tools that take a set of companies or a sector and return compact rows that cite passages, which keeps context small, cost predictable and answers grounded.

6. A portable graph layer

The same tool specifications compile to SPARQL for Amazon Neptune and to Cypher for Neo4j Aura, so the graph store is a deployment decision recorded in an architecture decision record, not a rewrite. Each request can name the store it queries.

7. Evaluation is a component, not a phase

A golden question set is written as paraphrase families: several wordings of the same question. The harness asks them as named test users and scores concept agreement (did different wordings resolve to the same meaning), evidence agreement (did they cite the same passages) and correctness. Extracted GAAP figures are scored against XBRL. A change to a prompt, model or tool is measured, not judged by feel.

8. Licensing enforced by construction

Data rights are modelled in the ontology itself, with an access level on every document. Licensed earnings call transcripts live in a separate private graph that is never loaded into Neptune, because a store whose default graph is the union of its named graphs would expose them to the public agent. The private view is reachable only with a scoped token, checked on every request. The boundary is structural, not a filter someone can forget to apply.

9. A public generative AI service with a known worst case

Opening an AI chat to the internet is a cost and abuse problem as much as a feature:

  • Scoped access tokens. Stored only as hashes, each in a pool with a daily dollar cap.
  • Reserve, then settle. Each question reserves its worst-case cost before any model call, and settles to the actual cost after, so the cap cannot be overrun.
  • Bounded agent loops. Each call is priced before it is made; steps and tool calls per step are capped.
  • Answer caching. A repeated opening question is served from cache at no charge.
  • Edge controls. CloudFront with origin access control and AWS WAF rules refuse floods before any compute runs; long answers are handed to a worker and polled, so no request outlives the edge timeout.
  • Guardrails. Amazon Bedrock Guardrails screen the question and the final answer.

10. Everything as code

The whole environment is Terraform, tested against a mocked provider, with significant choices recorded as architecture decision records. The extraction, the tools, the agent and this page are reproducible from the repository.

What carries over to the enterprise

Swap filings for contracts, policies, claims or engineering documents and the pattern holds: a governed vocabulary as the contract, verified extraction with provenance, a graph and an index joined on one identifier, agents that act with the user's own permissions, evaluation built in, and cost and data rights enforced by the architecture. The lab exists to show those pieces working together, with the trade-offs visible.

Designed and built by Dermot O'Brien, AI enterprise and solution architect.

All articles