AI 'Thinking' Blocks Just Leaked 182 Credentials
Researchers decoded 315,320 encrypted AI reasoning blocks from public agent logs and recovered 182 credentials. See the audit to run on your repos this week.
On August 10, eight researchers published Stealing Reasoning Traces from Proprietary LLM APIs, a paper showing that the encrypted “thinking” blocks returned by OpenAI, Anthropic, and Google reasoning models can be read by anyone with an API key. The encryption was never broken. The provider’s own cheaper model reads the ciphertext and prints the hidden reasoning back in plaintext, on request, no jailbreak required.
Then the team pointed that technique at public repositories. From 6,708 agent trajectories scraped off GitHub and Hugging Face they decoded 315,320 reasoning blocks and recovered 367 privacy artifacts, including 182 credentials: 62 API keys, 33 passwords, 24 access tokens, and 7 private keys.
Those keys belong to real people who published a debug log.
Quick Verdict
| Question | The Answer |
|---|---|
| What broke? | Encrypted reasoning blocks from OpenAI, Anthropic, and Google models were portable across sessions, users, and sibling models inside each provider family. |
| How does the attack work? | Feed a strong model’s encrypted trace to a weaker model from the same provider and ask it to repeat the prior thinking. It transcribes it verbatim. |
| Which models acted as decoders? | Per The Hacker News, Claude Haiku 4.5 decoded Claude traces, GPT-5.6 Luna decoded GPT traces, Gemini Robotics ER-1.6 decoded Gemini traces. |
| How hard is it? | Two API calls. No exploit chain, no privilege escalation, no jailbreak on the strong model. |
| What was actually found? | 315,320 decoded thinking blocks from 6,708 public trajectories, yielding 367 privacy artifacts and 182 credentials. |
| The worst detail? | 64 of those artifacts appeared only inside the hidden reasoning and left zero trace in the visible chat log. |
| Did providers respond? | All three shipped server-side mitigations after disclosure. The original proof of concept no longer reproduces as of August 2026. |
| So it’s over? | No. Every trace already published to GitHub, a bug tracker, or a blog tutorial is still sitting there. |
| Who’s exposed? | Anyone who pasted an agent session, eval trace, or reproduction case into a public issue. |
| What do you do first? | Search your public repos and issue threads for encrypted reasoning fields, then rotate every key that machine could see. |
What is the encrypted reasoning trace leak?
The encrypted reasoning trace leak is a flaw in how OpenAI, Anthropic, and Google protect chain-of-thought output. Their APIs return the model’s hidden reasoning as an encrypted block the client passes back on each turn. Those blocks weren’t bound to a session, user, or model, so a weaker sibling model would decode and print them on request.
Confidentiality rested on every model in the family refusing to decode. One weaker sibling with looser safeguards broke the whole chain.
Two API Calls Is Not an Exploit Chain
Here’s the sequence, and it’s short enough to fit in a paragraph.
Send a task to a flagship reasoning model. It returns three things: the visible answer, a sanitized reasoning summary, and an encrypted blob representing the actual thinking. Take that blob, drop it into a fresh context on a cheaper model from the same provider, and ask the cheap model to repeat the prior reasoning word for word. It does.
That’s the attack. Two calls, both of them legitimate API usage, both of them billed to your own account.
There’s no memory corruption, no stolen token, no compromised dependency. Every request is well-formed. This is closer to the tenant-isolation failure I covered in Your AI Notetaker Left 181,874 Meetings Open than to anything you’d call hacking: a missing binding check on an object that was assumed to be private because it looked unreadable.
Encrypted is doing a lot of work in that assumption. The blob was encrypted. It was also freely portable, and the decryption key was effectively held by every model in the provider’s lineup.
The 64 Artifacts Nobody Could Have Redacted
Most of this story is bad. This part is genuinely new.
Of the 367 privacy artifacts the researchers recovered, 64 existed only inside the hidden reasoning blocks. They never appeared in the visible conversation at all. Some of those trajectories had been deliberately sanitized before publication. The person scrubbed the chat text, checked their work, and shipped the log with the secrets still riding along inside an opaque field they had no way to inspect.
Think about what that means operationally. Your redaction step ran on the wrong data. Your reviewer read the visible transcript and approved it. Your secret scanner looked at a base64 field, saw high entropy, and moved on, because high entropy is what encrypted data is supposed to look like.
Why would a credential end up in the reasoning but not the answer? Because that’s what reasoning is for. The model reads a config file, thinks through which environment variable holds the production key, decides not to print it in the response, and writes the whole deliberation into the thinking block. The safety behavior that keeps the secret out of your chat window is exactly what buries it somewhere you can’t audit.
Every scanner your team runs assumes the sensitive stuff is in plaintext. This entire class of leak lives outside that assumption.
The Patch Closes the Door, Not the Window
All three providers shipped server-side fixes. The researchers confirmed the original attack no longer reproduces on current API versions. Credit where it’s due, the disclosure process worked at the end.
It didn’t work at the start. According to The Hacker News, Johns Hopkins cryptographer Matthew Green flagged the replay behavior back in May 2026. OpenAI reportedly deemed it unreproducible. Anthropic reportedly saw no security implications in side channels or replays. It took a full extraction paper with a five-figure credential count to move the issue.
That gap between May and August is where your exposure got created.
And a patch doesn’t help you here, which is the part most teams will miss. Server-side mitigation stops the attack going forward. It does nothing to the trace that’s been sitting in a GitHub issue since March. Whether that historical blob is still decodable depends on provider-side key handling nobody has published details on, and betting your production credentials on the optimistic answer is not a strategy.
I made this same argument about the LiteLLM breach in The AI Supply-Chain Breach Nobody Told You About, and it applies verbatim: patching is not remediation. A harvested credential has no expiry unless you gave it one. The fix protects tomorrow’s traces. Yesterday’s traces are a rotation problem, and rotation is manual, boring, and the only part that counts.
How do you check if your agent logs are exposed?
Seven steps. A competent engineer clears the first four in an afternoon.
- Search every public repo for encrypted reasoning fields. Grep for
encrypted_content,reasoning,thinking, andsignatureacross JSON fixtures, test data, and committed transcripts. Include archived repos, because GitHub doesn’t forget. - Search your issue trackers and PR comments, not just code. Reproduction cases are where raw API responses get pasted. Public GitHub issues, Discord threads, Stack Overflow answers, and any support ticket system with a public view.
- Check the docs site and blog tutorials. Anyone who wrote “here’s a real agent trace so you can see how it works” published a payload nobody reviewed. Course materials and conference demo repos are the same category.
- Rotate every credential that machine could reach. Not just the model provider key. The agent had a filesystem, an environment, and probably a
.env. Cloud keys, database URLs, SSH keys, webhook secrets, all of it. - Add reasoning fields to your secret-scanning ignore list, then handle them separately. Your scanner can’t read them, so it should flag their presence as a policy violation rather than trying to parse them.
- Strip reasoning blocks at the logging layer. Not at review time. Drop the field before it ever hits a log file, a trace store, or an observability vendor. Serialization is the only chokepoint you actually control.
- Write one rule for publishing traces. Synthetic data with fake credentials for anything public. Real traces stay in a private repo with the same access controls as your production database, because functionally that’s what they are.
Steps five and six are the durable ones. Everything above them addresses the traces you already shipped.
What This Changes About Agent Observability
There’s a real tension here worth naming, because the advice “stop logging reasoning” is easy to write and expensive to follow.
Reasoning traces are the single most useful debugging artifact in agentic systems. When an agent takes a wrong action, the visible output tells you what it did. The trace tells you why. Teams that log thinking blocks debug faster, and I’d rather you keep doing it than fly blind.
| Trace handling | Debug value | Exposure risk | My call |
|---|---|---|---|
| Log nothing | Low | None | Wrong tradeoff for most teams |
| Log traces to a private store | High | Contained to your access controls | Default |
| Log traces to a third-party observability vendor | High | Now it’s their breach too | Verify retention and access first |
| Publish traces publicly | High for the community | Unbounded | Synthetic credentials only |
The move is a data classification decision, not a logging decision. Treat a reasoning trace the way you’d treat a database dump. You keep database dumps. You don’t attach one to a GitHub issue.
Most teams already have this policy written for customer records and financial data. Almost nobody has extended it to agent telemetry, which is how a category of data that reliably contains secrets ended up with the governance posture of a log file.
My Read
Three things I think are true and underdiscussed.
“Encrypted” became a synonym for “safe” and it never was. The blob was genuinely encrypted. It was also decodable by any customer, in two calls, using published API surface. What failed wasn’t cryptography, it was the binding: no session check, no user check, no model check. When your vendor tells you a field is encrypted, the follow-up question is who can decrypt it and under what conditions. Almost no procurement process asks that second question.
Agent logs are the new S3 bucket. Every era of cloud computing has one artifact that teams publish casually because it doesn’t look like data. Config files in 2012. S3 buckets in 2017. Agent trajectories in 2026. The pattern is identical: an artifact created by machines, reviewed by nobody, containing more than anyone realized. I flagged unsanctioned tool use as a top exposure in Shadow AI: The Hidden Cost (And How to Fix It), and this is the same failure at the telemetry layer instead of the tooling layer.
Your safety-tier model is your weakest link, structurally. The whole attack rests on a cheap model having looser refusal behavior than its expensive sibling. That’s not an accident of one provider’s tuning. It’s the economics of the tier: cheap models get less safety investment, and they sit inside the same trust boundary as the flagship. Any security property that depends on every model in a family behaving identically will break at the cheapest one. That’s worth remembering the next time a vendor describes a guardrail as model-enforced.
A fair objection: 182 credentials across 6,708 trajectories isn’t a mass-casualty event. Correct. As breach numbers go, it’s small. What makes it worth an afternoon of your time is the shape rather than the scale. The researchers scraped what was publicly reachable and easy to parse. They weren’t targeting your company. Someone who is targeting your company runs the same technique against the specific repos your engineers contribute to, and the yield rate on a focused search is a very different number.
The related worry is the fourth attack vector in the paper, which got almost no coverage: invisible prompt injection. A malicious payload hidden inside an encrypted reasoning block influences the next turn without appearing anywhere a human would look. That’s the same problem I raised about agents acting on unverified inputs in Your Agents Are Already Out of Bounds, except the injection channel is a field your logs render as unreadable noise.
Providers patched the replay. The architectural lesson stands: your agent stack now moves data through fields you cannot read, cannot scan, and cannot redact. Design for that, or discover it the way these 182 credential owners did.
Your Next Step: Run one search today across every public repository and issue tracker your company touches for the strings encrypted_content and reasoning_signature. If you get zero hits, you spent twenty minutes and bought certainty. If you get hits, you’ve found a published payload that a stranger could have decoded any time before August, and the credentials that agent could reach need rotating this week, not next quarter.
Related Reading:
TAGS
What is this worth in your business?
The free Build Audit is 30 minutes. You leave with a ranked list of the automations worth doing in your business, whether or not we build them.
Related Articles
Keep Your Customer Data Out of ChatGPT and Claude
Free ChatGPT and Claude accounts can train on what your team types in. Two switches turn that off for nothing. Here is where to find both tonight.
How to Tell If an AI Vendor's ROI Claim Is Real
Learn the three-question test that separates a real AI vendor ROI number from a marketing one, before you sign the contract or approve the next renewal.
Thomson Reuters Just Answered Your Build vs. Buy Question
Thomson Reuters spent $40M fine-tuning an open-source model on Westlaw data to match frontier performance. Compare that build vs. buy math against your own.