How the ledger was read
This app reads a single cloud database of 4,183 published Australian public-sector accountability reports (2006–2026) from 50 oversight bodies across all nine jurisdictions — auditors-general, ombudsmen, anti-corruption and integrity commissions, royal commissions, the Productivity Commission, inspectors-general and independent reviews.
Each document was chunked and embedded (478,543 passages), mined for key points and 111,196 verbatim quotations, and independently graded by a language model on two 1–5 rubrics — finding severity and systemic extent — then tagged for theme and recommendation acceptance. A separate pass clustered the corpus bottom-up into 276 topics, rolled into 59 superclusters and 11 families, each re-graded across four eras and seven body types.
Everything on every page is a live query against Cloudflare D1 (the ~1 GB relational residual) and Vectorize (the embeddings), through fully-typed Drizzle. The same data is exposed to agents over an MCP server at /mcp.
How to read the output
These rules apply to every Insight Bridge corpus. They are properties of the method, not caveats about a particular run.
- Positions are relative to a proposition. A source’s position records how it stands against that cluster’s particular framing — not whether it agrees with some absolute claim. The same source can support one cluster and redirect a neighbouring one that covers similar ground differently.
- Propositions are synthesised from the cluster’s own members. Because the argument is built from the sources that were grouped together, and those sources are then assessed against it, a degree of agreement is built into the method. Comparisons between groups carry weight; a corpus-wide agreement rate does not.
- Counts describe the corpus, not the world. Every corpus here is curated. “N sources say X” measures what was collected and is never a measure of how common X is in the field.
- Clusters differ in how many distinct sources back them. A long document can fragment across many clusters, so weight a theme by the distinct sources beneath it rather than by how many clusters it contains.
- Every extraction and position is a model judgement. Key points, stances, propositions and syntheses are produced by a language model reading the source. They inherit its calibration and are not determinations of fact.
What the numbers are, in this corpus
Severity and systemic reach are graded on two 1–5 rubrics, so they rank reports against each other and never against an official standard. Cross-body divergence reflects both genuine differences of mandate and differences in what each body was examining, and averages combine very different report genres, so they describe the corpus rather than evaluate the bodies in it. Recommendation-acceptance rates cover only reports that recorded a response; many genres (royal commissions, Productivity Commission inquiries) structurally do not, and are excluded rather than read as evasion. The family tier is a clean is_primary partition, while the topic↔supercluster tier is genuine multi-membership. Partial-year 2026 is included as-is.