Key Takeaways
- SWE-bench Verified, a standard coding benchmark, rose from 60% to near 100% in a single year, while the Foundation Model Transparency Index average score fell from 58 to 40 over the same period, according to Stanford HAI's 2026 AI Index Report.
- Documented AI incidents rose from 233 in 2024 to 362 in 2025, a 55% increase, even as organizational AI adoption reached 88%.
- The most capable model developers, including OpenAI, Anthropic, and Google, have stopped disclosing training data composition, parameter counts, and training duration for several of their most resource-intensive systems.
- The gap matters operationally: the EU AI Act's Article 53 requires general-purpose AI model providers to give downstream deployers technical documentation, but a vendor disclosing less makes it harder for the deploying organization to complete its own risk assessments and compliance paperwork.
- Governance-specific hiring and policy adoption are both up, but organizations still cite knowledge gaps (59%), budget constraints (48%), and regulatory uncertainty (41%) as the leading barriers to closing the gap.
If your team just tried to fill out a vendor risk assessment for a frontier model and hit a wall of "information not disclosed," you are not imagining a trend. Stanford HAI's 2026 AI Index Report, released in April 2026, documents a widening split between what frontier AI models can do and what their makers are willing to say about how they work.
On one measure of coding capability, SWE-bench Verified, model performance climbed from 60% to near 100% in twelve months. Over the same period, the average score on the Foundation Model Transparency Index dropped from 58 to 40 out of 100. The two lines are moving in opposite directions, and the report's own framing is blunt about why that matters: "governance systems work best when the object of governance is visible," and AI is moving in the opposite direction, toward greater technical complexity and less public legibility.
For any organization that has to document a vendor's safety practices for an EU AI Act filing, a NIST AI RMF assessment, or an internal AI system registry, this is not an abstract research finding. It is a description of the paperwork getting harder to fill out, from the vendor side, at the exact moment more organizations are relying on that paperwork.
The capability side: benchmarks are being cleared faster than anyone expected
The 2026 Index's technical performance chapter is largely a story of benchmarks being solved. SWE-bench Verified, which tests whether a model can resolve real-world GitHub issues, went from 60% in the 2025 edition to near 100% in the 2026 edition. OSWorld, which tests whether an agent can complete real computer-use tasks such as processing documents or coordinating between applications, rose from roughly 12% to 66.3% success in two years, within six points of the human baseline. Stanford HAI's report also found that agents completing real-world tasks improved from a 20% success rate in 2025 to 77.3% in the current report, and that agents tackling cybersecurity capture-the-flag problems now succeed 93% of the time, up from 15% in 2024.
Adoption has kept pace with capability. Organizational AI adoption reached 88% in the 2026 survey data, and generative AI specifically is now used in at least one business function at 70% of organizations. The report also states that generative AI reached population-level adoption faster than the personal computer or the internet did, with consumer generative AI tools reaching an estimated 53% global population adoption within three years of ChatGPT's public launch.
None of that is in dispute, and none of it is slowing down. The concern the report raises is what is happening on the other side of the ledger while capability and adoption both accelerate.
The table below lines up the two trends the report tracks side by side. Read down each column separately first, then read across: the same twelve months that closed the capability gap widened the disclosure gap.
| Metric (Stanford HAI, 2026 AI Index) | 2024 / prior period | 2025 / current period | Direction |
|---|---|---|---|
| SWE-bench Verified (coding benchmark) | 60% | Near 100% | Capability up sharply |
| Foundation Model Transparency Index (avg. score) | 58 / 100 | 40 / 100 | Transparency down sharply |
| Organizational AI adoption | Not separately tracked at this level pre-2025 | 88% | Adoption up |
| Documented AI incidents (AI Incident Database) | 233 | 362 | Incidents up 55% |
| Businesses with no responsible AI policy | 24% | 11% | Policy adoption up |
| US private AI investment | N/A | $285.9B (vs. $12.4B in China) | Investment concentration up |
A rising number is not automatically good or bad here: incident counts and investment concentration are cautionary, while policy adoption is genuine progress. What is notable is that all of them are rising at once, alongside a falling transparency score.
The transparency side: the most capable models disclose the least
The Foundation Model Transparency Index, a project that scores model developers on how much they disclose about training data, compute, safety testing, and downstream impact, dropped to an average of 40 points in 2026 from 58 the year before. That reverses two consecutive years of improvement. According to the Index findings summarized in the report, training code, parameter counts, dataset sizes, and training duration have quietly stopped being disclosed for several of the most resource-intensive systems, including recent releases from OpenAI, Anthropic, and Google.
The distribution is uneven rather than uniformly bad. Some vendors, including IBM, scored close to the top of the scale, while others scored near the bottom, showing that low disclosure is a choice available to any vendor rather than an industry-wide technical constraint. The pattern the Index highlights is that the correlation runs the wrong way: the model families getting the most capability coverage in the same report are frequently the ones disclosing the least about how they were built and tested.
The report's authors describe this as a reversal, not a plateau. Two years of steady transparency gains across the industry gave way, in a single cycle, to a drop back below where the Index started. That timing overlaps almost exactly with the period in which frontier labs shifted toward closed release strategies for their most commercially significant models, which is consistent with the Index's own explanation for the decline.
The same pattern shows up in responsible AI benchmark reporting specifically. The report finds that almost all leading frontier developers report results on capability benchmarks, but reporting on responsible AI benchmarks (safety, fairness, and related dimensions) remains sparse by comparison. Among the model families the Index tracks, only one, Claude Opus 4.5, reports results on more than two of the responsible AI benchmarks the Index monitors. The report also flags a structural complication behind that gap: improving one responsible AI dimension can degrade another, so a developer optimizing a model for accuracy may see its safety or fairness scores move in the wrong direction, which makes the reporting itself more complicated to standardize even for developers willing to do it.
Why declining vendor transparency is now an operational governance problem
Here is where the finding stops being a research curiosity and starts being a compliance workload problem.
Under the EU AI Act, providers of general-purpose AI models are required, per Article 53, to maintain technical documentation and to make defined transparency information available to downstream providers and deployers who integrate the model into their own systems. That information must cover the model's capabilities, limitations, and how it was trained and tested, at a level sufficient for the deployer to understand the risks. If a foundation model vendor is disclosing less than it did two years ago, on the same categories the Transparency Index tracks, the deploying organization has less material to work with when it tries to complete its own side of that obligation: documenting the AI system in a registry, running a vendor risk assessment, or supporting a DPIA or AIA that depends on knowing what data a third-party model was trained on.
This is not a hypothetical mismatch. It is the same disclosure categories (training data, parameter counts, testing methodology) showing up on both sides. The Transparency Index tracks whether vendors disclose them. The Article 53 obligation requires vendors to disclose them to deployers specifically. When the first number falls, the second obligation gets harder to satisfy from the deployer's side, because the deployer cannot manufacture disclosure the vendor declined to provide. An organization running an AI vendor due diligence review this year is working with less vendor-supplied information than an organization running the same review in 2024, even though the regulatory expectation for what that review must cover has only gotten more specific in the meantime.
The incident data is the leading indicator, not a side note
The report also finds that documented AI incidents rose from 233 in 2024 to 362 in 2025, tracked through the AI Incident Database, a 55% year-over-year increase. Read next to the transparency numbers, this is the practical consequence rather than a separate story: incidents are rising in the same window that disclosure is falling and deployment is accelerating.
The report's own conclusion on this point does not hedge. It states that governance which is not keeping pace with deployment is already producing measurable harm, and that the pace is not slowing. That is a specific, falsifiable claim from Stanford HAI's own analysis, not an inference this article is adding on top of the data. It is worth reading exactly as written, because it changes the incident count from a curiosity into a trend line that governance teams should be tracking the same way they track deployment counts.
If your organization does not currently log AI-related incidents anywhere separate from general IT incident tracking, that is a gap worth closing before the next audit cycle, not after it.
What 88% adoption actually implies for governance maturity
An 88% organizational adoption rate does not mean 88% of organizations have a governance program sized to match. The Index's own responsible AI data shows the two numbers moving apart rather than together: AI-specific governance roles grew 17% in 2025, and the share of businesses with no responsible AI policy at all fell from 24% to 11%. Both are genuine progress, but the report frames this progress explicitly against the incident and transparency trends, describing responsible AI benchmarking as increasing but "not keeping up with AI advances and deployments."
The barriers organizations report are consistent with a governance function that is behind, not absent. Knowledge gaps (cited by 59% of respondents), budget constraints (48%), and regulatory uncertainty (41%) are the top three obstacles to responsible AI implementation. None of those is a lack of will; all three are capacity problems. An organization that adopted a generative AI tool in a single business function last quarter, and now has a governance team stretched across every other function's AI adoption too, is the median case this data describes, not an edge case.
What this means for vendor selection and AI system registries
Two practical decisions get harder when foundation model transparency drops, and both are decisions most organizations are actively making right now rather than at some future compliance deadline.
Vendor selection. A due diligence process built around "ask the vendor for their documentation" assumes the vendor will provide documentation as thorough as it did the last time you asked. The 2026 Index data says that assumption is now wrong for a meaningful share of frontier model providers. Vendor selection criteria should treat disclosure itself as a scored input: not just capability benchmarks, but whether the vendor publishes training data composition, safety testing methodology, and known limitations at a level that will actually hold up in your own documentation later. A vendor management process that centralizes vendor risk scores, certifications, and data processing agreements in one register makes this comparison possible across a growing AI vendor list, instead of re-litigating the same disclosure questions informally every time a new tool comes up for approval.
Internal AI system registries. Every AI system an organization deploys, whether built in-house or licensed from a vendor, needs an entry somewhere that records what it is, what data it touches, what risk tier it sits in, and who owns it. When the underlying vendor's own disclosure is thin, that registry entry becomes the organization's primary record of what it can verify about the system. Maintaining it well, rather than treating it as a one-time intake form, is the difference between a real risk picture and a stale one. A registry that maps each AI system to the regulations that actually apply to it, flags high-risk deployments automatically, and generates audit-ready documentation on demand does the work that vendor disclosure used to do more of.
Secure Privacy's AI Governance module is built around exactly this shortfall: it lets teams register every AI system with its risk tier, use case, and responsible owner, map each system to applicable regulations including the EU AI Act, and generate audit-ready documentation for regulators or internal oversight without depending entirely on how much the underlying model vendor chose to disclose. Paired with the platform's Vendor Management module, which tracks compliance status, certifications, and data processing agreements per vendor and flags compliance gaps automatically, it gives governance teams a way to keep their own records complete even when a vendor's transparency score is heading the wrong direction.
Common gaps this data exposes, and how to close each one
- No standing record of which AI systems are in use. If the honest answer to "list every AI system deployed across the company" takes more than a same-day meeting to produce, the organization is relying on institutional memory instead of a registry. Build one before a vendor's disclosure gets thinner, not after.
- Vendor documentation treated as a one-time intake step. A vendor that disclosed enough at signing eighteen months ago may not meet the same bar today. Re-request current documentation on a fixed schedule (annually at minimum, tied to contract renewal) rather than assuming the original file is still accurate.
- No internal owner assigned per AI system. The Index's own data shows governance roles growing but knowledge gaps still the top-cited barrier at 59%. An unowned system is the one that will not get re-reviewed when a vendor's transparency score drops.
- Governance maturity measured once, not tracked over time. A maturity score taken at a single point in time cannot show whether the gap between adoption and governance is widening or narrowing for your organization specifically, the same distinction the Index draws at the industry level.
Closing these gaps does not require waiting for a specific enforcement deadline. It requires a registry and a vendor management process that stay current by default, which is the operational fix this data is actually calling for.
FAQ
What is the Stanford AI Index and who publishes it?
The AI Index is an annual report published by Stanford University's Institute for Human-Centered AI (Stanford HAI). The 2026 edition is its seventh year and tracks technical performance, investment, adoption, policy, and responsible AI trends using data compiled from industry benchmarks, government statistics, and academic research rather than Stanford's own primary research.
What is the Foundation Model Transparency Index?
The Foundation Model Transparency Index is a research project that scores major AI model developers on how much they publicly disclose about training data, compute resources, safety testing, and deployment practices. Its average score across evaluated model developers fell from 58 to 40 in the period covered by the 2026 AI Index Report, reversing two prior years of improvement.
Does a lower Transparency Index score mean a vendor is breaking the law?
Not necessarily. Non-disclosure of training data or parameter counts is not automatically unlawful in most jurisdictions, though the EU AI Act does impose specific documentation obligations on providers of general-purpose AI models under Article 53. A low score is better read as an operational risk signal for deployers than as evidence of a legal violation on its own.
Why did documented AI incidents increase from 233 to 362?
Stanford HAI's report attributes the increase, tracked through the AI Incident Database, to the combination of faster AI deployment and governance capacity that has not scaled at the same rate. The report explicitly frames this as evidence that governance is not keeping pace with deployment, rather than as a data collection artifact.
How should a governance team respond to declining vendor transparency?
Treat vendor disclosure as a scored criterion in procurement rather than an assumed baseline, request the specific documentation categories that have declined most (training data composition, safety testing methodology, known limitations) before signing, and maintain an internal AI system registry detailed enough to stand on its own if a vendor's own documentation later proves thin.
Is generative AI adoption really faster than the PC or internet?
Stanford HAI's 2026 report states that consumer generative AI tools reached an estimated 53% global population adoption within three years of widespread public availability, which the report characterizes as faster population-level adoption than either the personal computer or the internet achieved in their first years. Organizational adoption, tracked separately, reached 88% in the same report.
The registry and vendor record are the part you control
Stanford's own data makes the diagnosis clear: capability is outrunning both transparency and governance capacity, and the resulting incident count is the receipt. None of that is something one governance team can fix industry-wide, and it does not need to be. What is fixable, directly and now, is whether your own AI system registry and vendor records are complete enough to withstand a vendor disclosing less next year than it does today.
Secure Privacy's AI Governance and Vendor Management modules exist for exactly that gap: a live registry of every AI system in use, mapped to the regulations that actually apply, backed by vendor risk records that flag compliance gaps automatically instead of waiting for the next audit to find them. See how the Privacy & AI Governance Platform handles AI system registration and vendor risk in one place, or book a demo to walk through it against your current regulatory footprint.




