Read me Page help ↗
Free Consultation
Observatory

Show your work

We argue everywhere else on this site that a governance claim should be checkable. This page is that standard turned on ourselves — what our agents changed without asking a human, what our own integrity agent says is wrong with us right now, every prediction this blog has committed to, and what we got wrong. Nothing here is curated for how it looks.

1
open findings against ourselves
none critical or high
no prediction settled yet
19 on the ledger · 1 past due
100%
of incident candidates rejected
684 judged, 0 kept
0
changes agents published with no human
0 in the last 7 days
2
official links we publish that are broken
out of 38 checked

The integrity agent last ran 4h ago. It runs every six hours against the live database and against this repository's own source files, and it has no way to mark itself healthy — it can only fail to find something.

What our own agent says is wrong with us

Open findings only, exactly as the agent filed them. Findings we have acknowledged or already fixed are not shown, because listing them here would pad this section with our own good news. The evidence payloads and the fix buttons stay in the admin dashboard; the statement of what is wrong is public.

low Data going stale open since 2026-09-06

2 regulation source check(s) could not complete

Timeouts, connection failures and server errors do not prove that a link is dead. Source watch retries automatically with backoff. Check the source and retry time below; investigate persistent failures. Acknowledge keeps the same condition in Reviewed; a new affected source or escalation reopens it.

The prediction ledger

Every original-thesis post this blog publishes has to commit to a dated prediction and say what would prove it wrong, or it does not get published. All of them are here, settled or not.

We have no hit rate yet. Nothing on this ledger has been settled, so there is no honest number to report — and a hit rate computed only over the ones we chose to settle would not be one either. 1 prediction is past due with no verdict, the oldest from 2026-06. Our own integrity agent raises a finding for each one until it is settled.

Past due, no verdict said it would be true by 2026-06

By June 2026, at least one major AI incident in an under-reported category (e.g., environmental harm or overreliance) will be revealed through a non-traditional source (e.g., academic paper, whistleblower, or citizen report) that the trade press had missed entirely.

The argument it came from: The concentration of AI incident reporting in just five legal and security trade outlets means that the AI risk landscape is being shaped by the legal profession's and security vendors' interests, not by actual system failures, and this is systematically hiding the most consequential risks like overreliance and environmental harm.

What would prove it wrong: No such incident emerges, and the trade press continues to be the primary source for all consequential AI incidents.

Read the post that made this call →
Open said it would be true by 2027-08

By 2027-08, at least one major agentic AI vendor will disclose a security incident where an API key compromise led to cascading access across multiple downstream systems, and this incident will be the first to be reported in both the security and governance news feeds simultaneously, marking the convergence of the two communities.

The argument it came from: The security community's top concern — 'Your AI's API Key Is a Master Key' and 'Supply Chain Attack' (6 articles) — points to a risk that governance frameworks do not address: AI systems inherit the access rights of the tools they invoke, so a single compromised API key in an agentic workflow can trigger cascading failures across multiple systems, and this is a governance gap that will be exploited before the 2028 EU AI Act embedded high-risk obligations take effect.

What would prove it wrong: If no such cascading API-key incident is disclosed by mid-2027, and the governance and security feeds remain disjoint on this topic, then either the risk is lower than the security community believes or the governance community is even more disconnected than argued.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, at least one major EU-based AI developer will publicly announce a reduction in generative AI features for EU users, citing the NCII/CSAM ban's compliance burden, and will relocate development to a non-EU jurisdiction.

The argument it came from: The EU AI Act's NCII/CSAM ban (effective 2026-12-02) will create a 'false positive crisis' for legitimate AI developers, because the ban's strict liability for generating prohibited content will force developers to over-filter outputs, and the incident data shows that 'genai' stories have dropped by 8 in the last 45 days—evidence that developers are already self-censoring, which will suppress legitimate use cases and drive innovation offshore.

What would prove it wrong: No such announcement occurs, and EU-based generative AI usage and development continues to grow at the same rate as before the ban.

Read the post that made this call →
Open said it would be true by 2027-12

Enterprise legal spend on AI compliance documentation will grow by over 50% while internal technical audits of AI models decrease.

The argument it came from: The divergence between compliance stories rising and governance stories plunging proves that enterprises are treating regulatory deadlines as paperwork exercises rather than security upgrades.

What would prove it wrong: If enterprise budget allocations show technical audit spending growing faster than legal documentation spending in annual risk surveys through 2027

Read the post that made this call →
Open said it would be true by 2027-12

By 2027-12-02, at least one major enterprise enforcement action under EU AI Act Annex III will target a multi-agent orchestration failure where no single vendor or deployer can be assigned sole accountability.

The argument it came from: Multi-agent systems will become the primary vector for untraceable corporate liability because current governance frameworks treat agents as single-instance tools rather than distributed supply chains.

What would prove it wrong: If all EU AI Act Annex III enforcement actions prior to this date name single-vendor, monolithic AI systems without multi-agent handoffs.

Read the post that made this call →
Open said it would be true by 2027-02

By 2027-02, the average salary for an AI model-risk validator in North America will have risen by at least 25%, and at least 20% of OSFI-regulated institutions will publicly warn of compliance delays.

The argument it came from: The OSFI Guideline E-23, effective 2027-05, will force Canadian financial institutions to treat AI models as regulated 'models' under existing model-risk management, and this will create a talent bottleneck that drives up compliance costs by 40% for early adopters, a cost that is not being budgeted for.

What would prove it wrong: If salaries rise less than 10% or no institution warns of delays, the talent bottleneck is overstated.

Read the post that made this call →
Open said it would be true by 2028-08

By 2028-08, the first major regulatory enforcement action under EU AI Act Annex III high-risk rules will target a robustness or overreliance failure rather than a data privacy leak.

The argument it came from: The absolute silence around model robustness and overreliance in incident feeds is a false-security artifact created by the fact that top-tier outlets prioritize privacy leaks over silent capability failures.

What would prove it wrong: If all EU AI Act Annex III enforcement actions prior to August 2028 are exclusively privacy or data governance violations.

Read the post that made this call →
Open said it would be true by 2027-01

By 2027-01, at least one major AI governance framework will add 'memory manipulation' or 'agent tool invocation' as a named risk category, directly borrowing from security research.

The argument it came from: The security community is 18 months ahead of the governance community on AI risk, and the gap is not a lag but a structural misalignment: security writes about actively exploited vulnerabilities and supply chain attacks, while governance writes about privacy and fairness, meaning the risk teams are defending against yesterday's threats while the attackers have already moved to agentic and memory-based attacks.

What would prove it wrong: If governance frameworks continue to ignore these attack vectors, the structural gap claim is overstated.

Read the post that made this call →
Open said it would be true by 2027-09

By Q3 2027, more than 50% of major AI security breaches reported in trade press will originate from compromised data science notebooks or unmanaged API keys rather than prompt injection or algorithmic bias.

The argument it came from: Security teams are tracking active exploitation of API keys and data science notebooks while governance frameworks are writing policies about fairness and bias.

What would prove it wrong: If prompt injection and algorithmic fairness issues account for the majority of reported enterprise AI breaches in public databases by Q3 2027.

Read the post that made this call →
Open said it would be true by 2027-04

By 2027-04-01, legal commentary and CPPA guidance will reveal that companies with large language models are claiming they are exempt from ADMT opt-out requirements by asserting 'human-in-the-loop review,' while companies with simpler automated scoring systems are being forced to comply—and the ratio of ADMT exemption claims for neural models vs. rule-based systems will be at least 3:1 in published compliance guidance.

The argument it came from: The 2027 California ADMT opt-out and pre-use notice requirements will create a perverse incentive for companies to deploy less transparent, less explainable AI models—because a model that cannot articulate why it made a decision is easier to claim is 'not subject to ADMT' than a rule-based model whose logic is auditable and therefore demonstrably automated.

What would prove it wrong: If the CPPA's final ADMT regulations explicitly state that any model that generates a decision affecting a consumer is subject to opt-out regardless of human review, and companies with black-box models are found to be complying at the same rate as those with transparent systems, this claim is wrong.

Read the post that made this call →
Open said it would be true by 2027-06

By 2027-06, at least one major AI governance framework will explicitly incorporate MITRE ATLAS techniques into its risk taxonomy, and public incident reporting will start citing these techniques.

The argument it came from: The attack techniques with documented real-world cases (e.g., LLM Prompt Crafting, 22 cases) are absent from the news window because the governance world is focused on privacy and fraud, not on the adversarial techniques that security teams are actively tracking—and this gap means governance frameworks are being built on the wrong threat model.

What would prove it wrong: If no major framework or regulator references ATLAS techniques by then, the misalignment is not being corrected.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, at least one major breach involving AI agent tool invocation (where an attacker, not a user, triggered an agent to exfiltrate data or transfer funds) will be disclosed in security press, and it will be initially misclassified as a 'multi-agent failure' or 'user error' in at least one governance-oriented publication.

The argument it came from: The MITRE ATLAS technique 'AI Agent Tool Invocation' (AML.T0053, 15 documented real cases) is the single most under-reported attack vector relative to its real-world prevalence, and the 70 stories on multi-agent risks in our feed are mostly about agent failures, not about attackers deliberately invoking agent tools to cause harm — meaning governance teams are preparing for accidents when they should be preparing for targeted abuse.

What would prove it wrong: If no such attack is disclosed by 2027-03, or if all disclosed agent-tool attacks are correctly attributed to external attackers in first reporting, then the framing gap is not as severe as claimed.

Read the post that made this call →
Open said it would be true by 2028-08

By 2028-08, when the EU AI Act Annex I embedded high-risk obligations apply, at least two major industrial AI deployments will face sudden regulatory halts due to unmeasured robustness failures that never appeared in prior incident feeds.

The argument it came from: The absence of reporting on environmental harm and system robustness is a structural failure of incident databases, not an indicator of zero risk.

What would prove it wrong: If all regulatory halts or enforcement actions under EU AI Act Annex I prior to August 2028 stem exclusively from risks already tracked in current top-five incident categories.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, public governance incident reports will continue to decline even as total internal AI incidents (as measured by corporate disclosures in securities filings) rise.

The argument it came from: The steep decline in governance-related AI incident reporting is not a sign of improvement but a leading indicator that organizations are shifting incident disclosure from public channels to private legal and compliance workflows ahead of enforceable deadlines.

What would prove it wrong: If public governance incident reports increase while compliance reports also rise, the shift hypothesis is wrong.

Read the post that made this call →
Open said it would be true by 2027-07

By July 2027, at least 40% of major enterprise AI security breaches will stem from robustness failures categorized as non-existent in current governance frameworks.

The argument it came from: Security intelligence briefings show 25 active ai-security articles while governance feeds report zero robustness stories, proving that technical exploitability is completely decoupled from enterprise risk registers.

What would prove it wrong: If enterprise AI incident disclosures attribute less than 10% of breaches to robustness or capability failures in published post-mortems through July 2027.

Read the post that made this call →
Open said it would be true by 2027-09

By 2027-09, a leaked internal document from a Fortune 500 company will show overreliance incidents were tracked internally but excluded from public incident reports, and this will trigger a class-action lawsuit.

The argument it came from: The 0 stories on 'Overreliance and unsafe use' (MIT AI Risk Repository 5.1) in 180 days is not a data gap but evidence that governance teams are actively suppressing this risk class because acknowledging it would invalidate their own AI literacy training programs, which are the primary compliance output they sell internally.

What would prove it wrong: If no such leak or lawsuit occurs by that date, and no regulator questions the absence of overreliance reporting, the prediction fails.

Read the post that made this call →
Open said it would be true by 2027-03

By early 2027, at least two major AI incidents will be revealed to have been reported as 'compliance issues' in 2026, with the underlying vulnerability still unpatched.

The argument it came from: The collapse in reported AI incident volume (from 450 to 344 stories) is not a decline in risk but a migration of incidents into the 'compliance' category, indicating that organizations are re-labeling failures to fit regulatory checklists rather than fixing underlying system flaws.

What would prove it wrong: If a retrospective audit of 2026 incident reports shows no systematic re-labeling from governance to compliance categories.

Read the post that made this call →
Open said it would be true by 2028-02

By February 2028, at least two of the Big Four accounting firms will formally combine their AI governance and cybersecurity risk practices into a single unified service line.

The argument it came from: Security intelligence briefings warning that 'Hackers want your AI brains more than your bitcoin' will force the merger of traditional CISO vulnerability management with model risk governance by early 2028.

What would prove it wrong: Public corporate restructurings showing continuous separation of AI governance and cybersecurity advisory practices among all Big Four firms through February 2028.

Read the post that made this call →
Open said it would be true by 2028-01

Publicly disclosed internal AI algorithm failures by Fortune 500 companies operating in California will decline by at least 25% in the twelve months following the January 2027 CPPA effective date.

The argument it came from: The impending enforcement of the California CPPA ADMT rules in early 2027 will cause enterprises to aggressively suppress internal AI incident logging to avoid creating discoverable compliance evidence.

What would prove it wrong: If public disclosures of internal AI algorithm failures by Fortune 500 companies increase or remain flat through January 2028.

Read the post that made this call →

What the agents changed with no human

When the official text behind a regulation moves, an agent reads it and drafts the new status. It only publishes if three independent readings agree on the confidence, on the status category, and on whether the change is material at all. Any dissent and nothing is published. Every change it did publish is below, with the before and after.

No agent has published a change on its own yet. This stays empty until three independent readings agree on one — which is the bar, and most candidates never clear it.

What we threw away

The promoter has judged 684 candidates from the live news feed and kept 0. It last judged something 4h ago, with 341 waiting.

100% of everything it looked at was rejected. That is the intended outcome — most AI news is commentary, not an incident — and the reasons are below.

  • 593 not actually an incident
  • 91 too thin to write up

Links we publish that are broken

Every regulation page links its official text. A watcher re-fetches all of them on a six-hour cycle, so a dead one is our broken promise and gets listed here rather than quietly sitting on the page. A further 4 hosts refuse automated checks entirely; those are verified by hand and are not counted as broken.

dead link last checked 6h ago

Voluntary Code of Conduct on Advanced Generative AI

https://ised-isde.canada.ca/site/ised/en/voluntary-code-conduct-responsible-development-and-management-advance

Request timed out after 20 s.

The page that publishes this link →

How to check any of this

  • Every regulation page carries the date its own official source was last read, and warns you when it has moved since we wrote the summary — start at the regulation reference.
  • Every curated incident says what it was, what failed, and the control that would have prevented it — the incident radar.
  • Every prediction above links the post that made it. The posts are dated and were not edited afterwards.
  • Every dated obligation we publish has its own page and a calendar feed you can subscribe to and check against the official source yourself.
  • Methodology & Scoring gives the actual formula behind every score and classification on this site, so you can recompute one rather than trust it.