Sthenos. Technologies
Original Research
Sthenos Technologies Original Research

The Trust Gap: The 2026 State of AI-Generated Code in Production

I work with enterprises each and every day, and what I can tell you is that AI already writes a large and growing share of the code that reaches production. What I cannot tell you, and what most engineering leaders cannot tell me, is whether anyone changed how that code gets verified, or who owns the risk when it fails. This study exists to put numbers on that, and it needs your answer.

Survey open, data collection in progress
Take the 4-minute survey

18 questions, about 4 minutes, anonymous. Every headline number gets published on this page, whichever way it cuts.

The thesis

We are accumulating verification debt and calling it productivity

I have seen this in law firms, start-ups, and large enterprises: AI now writes a meaningful and rising share of the code that reaches production, and the verification around it, the review, the testing, the accountability, looks about the same as it did before the AI showed up. Ask an engineering leader how much of last quarter's production code was AI-generated and you get a pause. Ask who owns the risk when that code fails and the pause gets longer. That pause is what this study measures.

I call it verification debt: the growing distance between how fast AI writes code and how well an organization can still verify, attribute, and stand behind it. This survey is the instrument built to measure that distance honestly. If the data comes back showing no gap, I will publish that. I have sat through too many delivery reviews to expect it.

By the numbers

About this survey

These figures describe the study itself. Nothing has been collected yet. Every statistic below is a question we are asking, none of them is an answer we are claiming.

18
questions, non-leading, with explicit "we don't track this" options
~4
minutes to complete, anonymous
5
core themes: adoption, verification, incidents, accountability, agentic reality
6
headline hypotheses the survey is built to test

Audience: software engineering leaders (VP and Head of Engineering, CTO, engineering managers, staff and principal engineers, SRE and platform). Sample size and field dates will be reported here at publication.

The six questions

Headline hypotheses the survey will test

Each one is stated as an open hypothesis. The blanks are what your responses fill in.

Shadow AI code

What share of engineering leaders cannot say how much of their production code was AI-generated?

The first measure of the gap: you cannot verify what you cannot see.

The verification gap

How many review AI-generated code the same as, or less than, human-written code, even after adopting it at scale?

Adoption at scale with unchanged scrutiny. That is the whole thesis in one line.

Incidents vs response

How many had an AI-related production incident or near-miss, and how many changed their process afterward?

The distance between what happened and what changed afterward is the debt made visible.

Accountability

How many have a named owner for AI-output risk, and for how many does nobody clearly own it?

Unowned risk is the version of this that ends up in front of a board.

The agentic reality gap

Everyone is talking agentic. How many actually run agentic AI in production?

Separating the conference-stage roadmap from what actually runs in production.

Confidence vs proof

Is confidence shipping AI code rising faster than the verification meant to justify it?

Confidence without the verification to back it is debt with a smile on it.

Inside the instrument

What each section measures

The full survey runs five thematic sections plus demographics. Here is what each one is built to surface, and why I put it in.

A

How deep is AI in the codebase

The baseline first. How much of the production codebase is AI-assisted or AI-generated, how fast is that share moving, and which tools are actually driving it? You cannot manage a number nobody tracks.

  • Roughly what share of your team's production code was AI-assisted or AI-generated in the last 3 months?
    0% / 1 to 10% / 11 to 25% / 26 to 50% / 51 to 75% / 76%+ / We don't track this
  • How has that changed versus 12 months ago?
    Much lower / Lower / About the same / Higher / Much higher / We don't track this
  • Which AI coding tools are in regular use?
    GitHub Copilot / Cursor / Claude Code / Windsurf / Tabnine / Amazon Q / JetBrains AI / In-house / Other / None
Takeaway we are testing: my bet is a large share of leaders pick "we don't track this". That answer is itself the finding.
B

Verification, the gap

Whether scrutiny scaled with volume. AI can produce in an afternoon what a team used to ship in a week, so the question becomes whether the review, testing, and QA around it grew to match, or stayed exactly where it was.

  • Compared to human-written code, AI-generated code in your org gets:
    More review / The same review / Less review / No different process / Depends
  • Can you quantify how much AI-generated code reaches production without independent human review?
    Yes, we measure it / We estimate it / No, we don't track it
  • Have you changed testing, review, or QA specifically because of AI-generated code?
    Yes, significantly / Yes, minor changes / No, same as before / Planned but not done
  • Over the last 12 months, your verification investment relative to code volume has:
    Grown faster than volume / Kept pace / Fallen behind volume / Unchanged / Don't track
Takeaway we are testing: whether verification investment has fallen behind the volume of code it is supposed to check. I expect it has, and by more than most leaders would guess.
C

Incidents and accountability

Cause meets consequence here. Has AI-generated code already broken production somewhere, who owned it when it did, and could anyone even identify the AI-authored code after the fact? If you sell into financial services, healthcare, or government, that last question is not academic, an auditor will eventually ask it.

  • In the last 12 months, has your org had a production incident or near-miss traced to AI-generated code?
    Yes, incident(s) / Yes, near-miss only / No / Not sure / We could not tell
  • Who owns the risk when AI-generated code causes a production problem?
    Individual engineer / Eng lead / QA or test / A dedicated role / Nobody clearly / Not defined
  • How confident are you that you could identify which production code was AI-authored if an auditor asked?
    Very / Somewhat / Not very / Not at all
Takeaway we are testing: that "not sure" and "nobody clearly" come back far more often than they should on questions a mature organization can answer cold.
D

Agentic reality

Every roadmap conversation I sit in right now is agentic. This section separates that conversation from actual production deployment, and asks for the single biggest blocker keeping autonomous agents out.

  • Where is agentic AI (autonomous multi-step agents) in your org today?
    In production / Piloting / Experimenting / Evaluating / Not on the roadmap
  • The biggest blocker to trusting agents in production is:
    Reliability / Observability / Security / Cost / Verification and testing / Governance and compliance / We don't use them
Takeaway we are testing: that far fewer organizations run agents in production than the discourse implies, and that verification sits at or near the top of the blocker list.
E

Leadership and outlook

Takes the gap up a level. Does the board even see AI-code risk, and where is shipping confidence heading? There is also one open-text line for the biggest unsolved problem in getting AI code safely to production, because the best answers never fit a multiple choice.

  • How visible is AI-code risk to your non-engineering leadership or board?
    Actively tracked / Discussed / Aware but not tracked / Not on their radar
  • Your confidence shipping AI-assisted code to production today versus 12 months ago:
    Much lower / Lower / Same / Higher / Much higher
  • In one line, the biggest unsolved problem in getting AI-assisted code safely to production:
    Open text
Takeaway we are testing: whether confidence keeps climbing while board-level visibility of the risk stands still.
Definition

Verification debt, defined

Verification debt is the accumulating gap between how quickly an organization produces code and how well it can still verify, attribute, and stand behind that code. It behaves like technical debt: invisible on the day you take it on, expensive on the day it comes due.

AI raises the interest rate on all of it. Work that used to take a big team months can now be built in a fraction of the time, I have seen this in law firms, start-ups, and large enterprises, and the volume of output that needs checking grows far faster than the review, testing, and ownership around it. Someone still has to verify and supervise what the AI produces. Skip that, and you are borrowing against your own reliability one merge at a time. The trust gap is simply the day the debt comes due: the moment a leader cannot say how much AI code shipped, who reviewed it, or who owns it if it breaks.

Methodology

How the study works

  • Design. A primary, original survey of software engineering leaders. Nothing here is a synthesis of public data; every finding will come from self-reported responses collected on this instrument.
  • Instrument. 18 questions across five thematic sections plus demographics, about 4 minutes to complete, anonymous. I wrote the questions to be non-leading, and every tracking question carries an explicit "we don't track this" option, because the absence of measurement is data and I want it captured, not hidden.
  • Audience. VP and Head of Engineering, CTO, engineering managers, staff and principal engineers, and SRE and platform roles, across company sizes and industries including financial services, healthcare, government and public sector, technology and SaaS, and retail.
  • Reporting. Sample size and field dates get published on this page alongside the results. Where the data does not support a hypothesis, we say so. No number on this page is a finding until the results land here.

Note on sourcing: this page cites no third-party statistics because it reports primary research that has not been collected yet. Every figure shown describes the survey instrument itself.

Add your data to the record

This report is only as strong as the honesty of the people answering it. Four minutes, anonymous, and the industry gets a real benchmark on how AI-generated code actually gets verified, instead of one more opinion about it.

Take the 4-minute survey
Cite this report

Citation

Sthenos Technologies. "The Trust Gap: The 2026 State of AI-Generated Code in Production." Original research, July 13, 2026. https://sthenostechnologies.com/ai-to-production-gap-2026/

Journalists and analysts are welcome to reference the framing now, and the findings once they publish. If you want an early read of the results, the data cut by role, company size, or industry, or an on-the-record comment from the author, contact Sthenos Technologies.

SS

Suleman Siddiqui

Advisor to the CEO at Sthenos Technologies

Suleman Siddiqui leads strategy at Sthenos Technologies and works with enterprises each and every day on projects large and small, from law firms to public-sector systems. This report came out of a pattern he kept running into across delivery teams: AI writing more and more of the code, and almost nothing changing about how that code gets verified. The survey exists to find out whether that pattern is as widespread as it looks from inside the delivery work.

Wrestling with AI-generated code in your own pipeline?

If verification debt is a live problem on your team, I am happy to compare notes. No pitch, just a conversation about what is working and what is not.

Book a 30-minute call