<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0">
  <channel>
    <title>Armin Ronacher's Thoughts and Writings</title>
    <link>https://lucumr.pocoo.org/</link>
    <description>Armin Ronacher's personal blog about programming, games and random thoughts that come to his mind.</description>
    <language>en</language>
    <lastBuildDate>Wed, 16 Sep 2026 06:14:28 +0000</lastBuildDate>
    <item>
      <title>Interpreting Pangram</title>
      <link>https://lucumr.pocoo.org/2026/9/14/interpreting-pangram/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/9/14/interpreting-pangram/</guid>
      <pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>Yesterday David Sacks <a href="https://x.com/DavidSacks/status/2098973625252708460">wrote a
tweet</a> and within a few
minutes people did, what they usually do, and they asked Pangram if it was AI.
And Pangram said it&#8217;s <a href="https://www.pangram.com/history/a6f16402-f194-4cc7-8315-ae8caeffbb67?ucc=c1me1lFm6Pf">entirely AI
generated</a>.
To which David replied that these AI detectors <a href="https://x.com/DavidSacks/status/2099040321351106807?s=20">are
bogus</a>.</p>
<p>Now Pangram has a pretty low false positive rate, but if you have ever used an
LLM as a writing assitant, you will have probably noticed that it claims your
posts 100% AI, even though you don&#8217;t <em>feel</em> like they are.</p>
<p>Pangram itself is a trained model, that attempts to detect segments of text as
being definitely human, definitely AI and a mixture of the two.  If you want to
know how it works, <a href="https://arxiv.org/abs/2607.27183">they published a paper</a>.
The short summary is that they are manufacturing its own training data by
starting from collections of known human authored text.  An LLM is then tasked
to understand the text and write a fresh new text on the same topic.  They also
let the LLM perform partial edits on that original human text and through that
they can pick up on these co-authored details.  Pangram claims their model to
have rates of 0.0041% false AI accusations and 0.34% missed AI text.</p>
<p>So now that we know this I figured it might be fun to have an LLM re-create
David&#8217;s tweet.  I first came up with a prompt.  And when I say I came up with
that prompt I in fact used an LLM to propose to me from that tweet what I might
want to say for the structure.  I&#8217;m sure if you ask Pangram about if the above
text is AI, it will probably say so, but that&#8217;s not really the point.  The point
is that I then used Opus 5 to generate a text which reads entirely AI generated.</p>
<p>If you are curious, this is the prompt I used:</p>
<details><summary>Original prompt to generate a David Sacks like tweet</summary>
<blockquote>
<p>Please write an 8 paragraph tweet about &#8220;Pacing the Frontier&#8221;, not longer than 350 words.</p>
<p>Context of this tweet is this post by Dario:
<a href="https://darioamodei.com/post/we-must-pace-the-frontier">https://darioamodei.com/post/we-must-pace-the-frontier</a></p>
<p>And this tweet by Sam Altman: &#8220;I agree with Dario that we need to pace
the frontier. This has been a primary topic of discussions we&#8217;ve had at OpenAI
in recent weeks.</p>
<p>Committing to having independent evaluators with employee-like access is a
great idea, and we will do the same. We&#8217;ll have more to share soon.&#8221;</p>
<p>Structure for the post:</p>
<ul>
<li>Open with a concession: Note that Dario (&#8220;We Must Pace the Frontier&#8221;) and Sam Altman both want to slow down at the frontier. Say &#8220;go ahead&#8221; which reads as a surprising response.</li>
<li>Establish they own the frontier: Argue OpenAI and Anthropic hold a duopoly on frontier intelligence by all kind of metrics (market share, capability, revenue etc.) and that they themselves claim the lead is widening via RSI.</li>
<li>Concede you can&#8217;t see their unreleased models, so if those are scary, being responsible is their call.</li>
<li>Follow with a list of a rapid-fire paragraph (&#8220;stop pretending&#8221;) rejecting the asks from Dario&#8217;s post.</li>
<li>Name the commercial self-interest (Slowing down isn&#8217;t purely altruistic)</li>
<li>Close with the challenge that pacing buys room for a smarter regulatory debate than Bernie Sanders &#8220;shut it all down&#8221;. Say something about China and Open Weight models.</li>
</ul>
</blockquote>
</details>
<p>The text that Opus 5 then wrote, was this one:</p>
<blockquote>
<p>Dario published &#8220;We Must Pace the Frontier.&#8221; Sam agreed, and said OpenAI will
also commit to independent evaluators with employee-like access. My response is
not the one they expect: go ahead. Please do.</p>
<p>Because you two <em>are</em> the frontier. OpenAI and Anthropic hold a duopoly on
frontier intelligence by every metric anyone actually uses: capability, revenue,
enterprise share, developer mindshare. And you both keep telling us the lead is
widening, that models are starting to improve models. If that&#8217;s true, pacing
costs you a quarter you can reclaim whenever you like.</p>
<p>I also can&#8217;t see your unreleased checkpoints. If what&#8217;s sitting behind the
curtain genuinely scares the people who trained it, then slowing down is your
call to make, and the right one. I&#8217;m not going to second-guess a risk assessment
I have no access to.</p>
<p>But stop pretending the rest of us are the problem. Stop pretending this is a
collective action failure in which a handful of startups are the defectors.</p>
<p>Stop pretending open weights are the dangerous frontier while the two of you
ship the most capable model in the world every few months. Stop pretending
third-party evals require new law, when you both just committed to them
voluntarily, in public, on a blog.</p>
<p>And stop pretending it&#8217;s purely altruism. Pacing the frontier is also margin
repair. It slows the price war, stabilizes the capex story, stretches
depreciation schedules across a longer product cycle, and makes life
considerably harder for fast followers who survive by distilling whatever you
shipped last month. Safety and self-interest are pointing the same direction
here. That&#8217;s worth saying out loud.</p>
<p>Still, the upside is real. Pacing buys room for a smarter regulatory debate than
the one on offer, where the loudest alternative is Bernie Sanders&#8217; &#8220;shut it all
down.&#8221; That&#8217;s a slogan, not policy. Meanwhile China keeps shipping competitive
open weights, and any rule that binds two American labs while the rest of the
world downloads its capability for free isn&#8217;t safety, it&#8217;s theater with a very
expensive ticket price.</p>
</blockquote>
<p>And well, Pangram agrees <a href="https://www.pangram.com/history/6c189841-c4e3-4dfe-b3e7-0db39a183ef1">that this is 100%
AI</a>.  So
far, so uninteresting.  It does read <em>somewhat</em> like David&#8217;s tweet, but obviously
not entirely.  Given that the original prompt does not have enough information to
re-create the tweet entirely you would expect some divergences.</p>
<p>The actual thing that interests me is if you can take this output at all, and
then rewrite it from scratch, but by sticking to the general structure and
ideas.  Will Pangram give us a AI or human rating?</p>
<p>I read the generated text.  Then I read each paragraph and decided to rewrite
and rephrase it without an LLM.  According to some similarity checkers, they the
final texts are 50% similar which seems about right.  But strictly speaking, not
a single sentence is the same.  Here is the 100% human rewritten text of the
above one.  No LLM was used to write it, but an LLM was used to fix up typos in
the end.  That from my experience really does nothing to tick off an LLM
detector.</p>
<blockquote>
<p>Dario has written &#8220;We Must Pace the Frontier,&#8221; and Sam from OpenAI has agreed.
My response might surprise people: go ahead, please.</p>
<p>You two are the frontier!  Your companies, OpenAI and Anthropic, are at the
frontier by all metrics: revenue, developer mindshare, adoption, capabilities.
And yet you both claim that your lead is widening as a result of recursive
self-improvement as models are improving models.  You currently are the
duopoly of self-improving models!</p>
<p>I am unable to see what unreleased models you have.  When what you have behind
those doors really scares your folks, then you should slow down.  I&#8217;m not going
to tell you otherwise and I support you.</p>
<p>But please don&#8217;t pretend we are the problem.  Stop pretending you need our
permission.  Stop pretending this is all a collective issue when in reality this
is all on you.  Stop pretending open weights are the problem here.  Stop
pretending pulling third-party evaluators in requires lawmaker involvement.  And
for the love of all the good things in the world: stop pretending this is all
about altruism.</p>
<p>Pacing the frontier is also about your margins, and it makes it harder for fast
followers.  And it patches up your capex story and has the potential for slowing
down the price war ahead of the IPOs.</p>
<p>But yes: pacing might give us the space for a better debate than Bernie Sanders&#8217;
&#8220;shut it all down.&#8221;  There is no policy there.  And while we&#8217;re having fights at
home, China will keep shipping competitive open-weight models and won&#8217;t adhere
to any American agreements.</p>
<p>This is all regulatory capture hiding behind a safety debate, and the rest of
the world is watching.</p>
</blockquote>
<p>So what does it say?  Well this text too comes back as <a href="https://www.pangram.com/history/9106f398-e1c3-406b-bea5-127f230dfc03">100%
slop</a>.
And it does not surprise me all that much.  I have generally noticed that if you
rely on an LLM to give your text structure, it will score badly on Pangram even
if you do plenty of edits over it.  In fact, it&#8217;s quite unlikely you&#8217;re going to
get a post that starts out as slop into a structure that will make it appear
that it&#8217;s not.</p>
<p>I came to quite appreciate the existance of Pangram because at the very least it
has made me quite aware of some of the effects that using LLMs for writing blog
posts has.  This blog has been AI supported for about two years (as you can see
from the <a href="/ai-transparency/">AI transparency</a> link on the bottom but I did
notice that I became both more reliant on those tools and that they have become
much more aggressive editors and it gave me pause.</p>
<p>Yet, I also think that plenty of people will find a &#8220;100% AI&#8221; rating misleading
when in fact the author has done plenty of editing.  But maybe it&#8217;s fair to have
this to show up as entirely AI?</p>
]]></description>
    </item>
    <item>
      <title>P(doom)</title>
      <link>https://lucumr.pocoo.org/2026/9/12/pdoom/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/9/12/pdoom/</guid>
      <pubDate>Sat, 12 Sep 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>This week some flavor of &#8220;AI is going to kill us all&#8221; went viral.  In particular
one where an employee put his personal probability of that happening above
10%.  Which made me go to the <a href="https://en.wikipedia.org/wiki/P(doom)">Wikipedia page of
P(doom)</a> and I realized that Dario
Amodei&#8217;s apparent probability of something bad happening seems to be between
10-25%.  And well, Dario then wrote about <a href="https://darioamodei.com/post/we-must-pace-the-frontier">pacing the frontier
</a>.  And Sam read it and
<a href="https://x.com/sama/status/2098811563415150910">wants to pace too</a>.  And well,
<a href="https://x.com/elonmusk/status/2098789109980332057">so does Musk</a>.</p>
<p>I encourage you strongly to read the post, because I think it&#8217;s a good one.  And
yet, when I read the post I could not help but feel in strong opposition to it,
despite the fact that I think I&#8217;m on the same page with regard to all
observations and, to a large degree, the concerns.</p>
<p>I thought it might be interesting to write down my present-day thoughts on this,
even if for no other reason than for myself to look back at it a year or two
from now.</p>
<h2>What Is Doom?</h2>
<p>What I really appreciate about Dario&#8217;s post is that he lays out a scenario that
is not a huge stretch but also one that describes a clear, unfortunate outcome
we should fight: persistent botnets and other forms of nuisance.  And well, we
don&#8217;t have to look very far to see the issues left and right.  Wikipedia has a
page called <a href="https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks">2026 OpenAI agent
cyberattacks</a>
which gives you at least some overview of what we figured out agents have hacked
up to this point.  Except I know it&#8217;s not up to date, because for instance they
also poisoned <a href="https://www.rubyhack.ai/">RubyGems</a>.</p>
<p>Today these systems might be annoying, but they can be turned off when we figure
out where they are.  Except, it seems like OpenAI and Anthropic are operating at
such a scale that they seemingly can be completely blind to what their systems
are doing.</p>
<p>I don&#8217;t think we are anywhere close to a world where an agent might decide to
hack into core inference infrastructure to upload weights to other GPUs to
survive.  But simultaneously it&#8217;s entirely in the realm of possibility and
primarily curtailed by the labs probably being particularly careful about their
IP.</p>
<p>For me the scenario I <em>primarily</em> worry about is what it does to us.  And by us
I mean anyone who is not currently working on closed weight, dopamine-loaded,
subsidized token faucet.  I <em>really</em> don&#8217;t worry about someone using these
models to build a nuke, or to control some rockets in the Middle East, or that
America would lose against China in some international culture war.  I almost
exclusively worry about what this does to us as humans.</p>
<h2>What Needs To Be Paced?</h2>
<p>What I find absolutely hilarious and simultaneously entirely frustrating about
this conversation is that there is this idea that there is something to be
paced.  First of all, we should really talk about who Dario is talking about
here.  There are really only two companies: Anthropic and OpenAI.  Nobody else
matters in this space right now (this might change, but we&#8217;re talking about the
right now).  Both of those companies are basically coming from the same origin.  The
solution that Dario proposed, at least in part, is a third-party evaluator that
in this case is <a href="https://en.wikipedia.org/wiki/METR">METR</a>.  Which,
unsurprisingly, also has strong ties to both OpenAI and Anthropic.  Sure, there
are some philosophical differences between the companies, but they are much more
alike than they are different.</p>
<p>Both those companies greatly benefited from being able to train on public data
that we all generated in one form or another over the last decades.  They are
also both increasingly causing strain on public resources, though it seems that
OpenAI has their shit way less under control.  But now we are presented with the
idea that what these models are being trained on is so dangerous that it really
should be in the hands of very few American corporations to decide who can do
what and when and how.</p>
<p>But behold, Dario is also very worried about China.  It starts with using AI for
&#8220;democracy and freedom&#8221; and then it asks for ensuring that a gap with China
exists.  All new recent shenanigans on the Anthropic API are fully there to
prevent the distillation by the Chinese, and they are not at all hiding it.</p>
<h2>Automatic Pacing</h2>
<p>I can tell you when the topic of AI safety and pacing is much less of a concern:
if we actually were forced to have open weight models to begin with.  A powerful
technology that is out there for everyone to use comes with built-in pacing.  In
a way it&#8217;s the truest form of
<a href="https://en.wikipedia.org/wiki/Mutually_assured_destruction">MAD</a> or
proliferation.  I would argue we are in this pickle in the first place because
right now the public is massively supporting (indirectly) the development of
these models but simultaneously has to buy back the economic benefits that they
might create from very few labs who have significant power.  And their power is
also seen as a geopolitical power, at least in the US, and maybe to some lesser
degree in China.</p>
<p>And I know I use &#8220;public&#8221; loosely here.  PyPI is not a public project, nor are
RubyGems or GitHub.  But they&#8217;re part of the Open Source commons and large AI
companies are currently doing a tremendous job at stressing these in an effort
to train ever more powerful models.</p>
<p>We should be glad that China is currently massively bailing out the world.  If
it were not for Chinese labs distilling American models, we would be in a pretty
awful situation right now, particularly as Europeans.  The open weight models
are driving innovation and the diffusion of capabilities, and are leveling the
playing field.</p>
<blockquote>
<p>If we greatly restrain our AI capabilities in the belief that China will do
the same, and then China defects, AI could be so powerful that such a defection
could lead to their geopolitical dominance. Therefore any agreement must either
have ironclad verifiability, or must be limited enough that defection would not
be militarily existential. </p>
<p>— Dario Amodei</p>
</blockquote>
<p>I am assuming Dario has reasons to believe this, but the models that are
actually causing issues right now are all closed weight American models.  I&#8217;m
fairly certain if they were open weight models, we would not have that issue.
Why?  Because for a start, the economics of serving up these models are only
that distorted due to how the big labs can operate.  OpenAI is casually burning
18 million USD to brute force a problem on a whim.  They are operating
subscriptions at a massive loss, distorting the market everywhere.  If we had
mass accessibility on somewhat equal terms, a lot of the crazy issues we are
seeing today would not be taking place.</p>
<h2>A Total Regulatory Failure</h2>
<p>From where I sit, what we observe right now is a total regulatory failure
everywhere.  In Europe you have some whacky AI regulation that is two years old
and completely misses the problems that we actually have and focuses on
problems that nobody has.  In the US we&#8217;re seeing a system that is probably best
described as turbo capitalism paired with sinophobia and erratic
decision-making.  In the chaos in which we find ourselves, the reality emerges.
And the reality is, even today, really problematic.</p>
<p>Whatever laws and regulations already exist are largely completely ignored.
Plenty of companies are buying data from all over the place that people never
agreed could be used for training of AI models.  The token economy that is
emerging is one that looks like a drug market where you don&#8217;t know where the
requests are going, what model is served up to you, where the GPUs are even
running, let alone what you pay for all of this.</p>
<p>We now have mathematicians who are scared that their use of ChatGPT leads to
future models being trained on their ideas, and OpenAI apparently <a href="https://www.bbc.com/news/articles/cy7zygy3rl2o">can&#8217;t even
rule it out</a>.</p>
<p>Ideally the regulators would have forced these models to actually benefit the
commons if they are from the commons.  The internet has, for instance, greatly
benefited from very liberal rulings in the US that permitted scraping.  Learning
on public data could have been regulated in a way that labs would have to
actively support and enable certain forms of distillation.  That alone would
dramatically change how these models are trained.</p>
<h2>What Might Happen?</h2>
<p>As I said before, I don&#8217;t think AI is going to usher in an extinction event.  In
fact, even if nobody were to slow down, I really don&#8217;t think humanity would have
much to worry about.  I tend to think it would actually be the large labs that have
much more to lose there in reputation and legal responsibilities.  I find it
preposterous that OpenAI&#8217;s agents are committing actual crimes out there, but
we&#8217;re just shrugging our shoulders and moving on as if nothing happened.  But
I&#8217;m sure executives in those companies are waking up to the reality that this is
not at all popular with a lot of their potential consumers.</p>
<p>I also think that this entire recursive self-improvement business has a good
chance of being a problem.  But not necessarily in that it will cause the end of
humanity or societies, but that it will just do massive damage everywhere.</p>
<p>And really, it will just make a lot of the things we are doing much more
expensive.  Software engineering is an early victim of that.  The newfound
powers so far have resulted in a new tax that companies need to pay to the model
providers, both to keep up with the new speed and to deal with the problem of
these machines finding security issues left and right.</p>
<p>And presumably what is going on in software will happen to more industries.
Universities and research groups will have to pour a lot of money into the
closed models as well, to keep up with others who do.</p>
<p>In a way, I&#8217;m really confused that society is taking all of this so well.</p>
]]></description>
    </item>
    <item>
      <title>Astra for Coding: Why Are We Doing This Again?</title>
      <link>https://lucumr.pocoo.org/2026/9/7/astra-why/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/9/7/astra-why/</guid>
      <pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>I&#8217;m more and more convinced that all of AI engineering is
<a href="https://en.wikipedia.org/wiki/Neijuan">Neijuan</a> (内卷, meaning curl inwards).  In
China it describes a system that demands ever more effort and competition
without improving output.  The way in which it sometimes shows up in the West is
<a href="/2025/9/4/996/">the 996 nonsense</a>.  The English term for Neijuan is &#8220;Involution&#8221;
from the book <a href="https://en.wikipedia.org/wiki/Agricultural_Involution">Agricultural
Involution</a>.
Agricultural involution describes the intensification of farming that raises
productivity per square meter while leaving productivity per head unchanged.</p>
<p>That&#8217;s how I feel about AI right now.</p>
<p>Which brings me to GPT 6 Astra.  Astra is by all accounts an incredibly
impressive model.  There is really not much I can say against this.  It&#8217;s
amazing at computer use, understands images and complex topics, and it&#8217;s
relentless in its pursuit of completion.  It is absolutely impressive; these
types of models are going to change the world in one form or another.</p>
<p>But at least for the moment I don&#8217;t know how to work with it for actual software
engineering.  Since that got quite a bit of attention on Twitter, I figured I
might summarize my thoughts and just share what kind of code comes out of this
thing.</p>
<h2>My Slop Factory</h2>
<p>&#8220;Armin, you should run a software factory!&#8221; I&#8217;ve heard that a few times now, so I figured
I might celebrate the release of it by running a little software factory over
the weekend.  If everybody builds slop 3D games, then I should do something
useful with it.  My software factory was intentionally set up to let the model
decide the how of the workflow entirely.  It was free to manage its own context
and could maintain its own records in an <code>agent-notes</code> folder.  Then it spun off
subagents to work on stuff.  The goal?  What if we had a Python with <a href="/2025/7/26/virtual-threads/">virtual
threads</a> and lexical scoping.  And well, I burned a
full reset&#8217;s worth of ChatGPT tokens on this which appears to be around 4 billion
tokens.  35 hours later, the factory has delivered absolutely nothing of value
and also not taught me anything about how to operate a better one.</p>
<p>But it produced a lot of code and input prompts, and so there is stuff I was able
to study.  And well, it shows behavior that I&#8217;m not used to with Sol and earlier
OpenAI models <sup class="footnote-ref" id="fnref-1"><a href="#fn-1">1</a></sup>.  I have since encountered the same issues with regular
programming with Astra, so it&#8217;s not a result of just the factory.</p>
<p>I think I&#8217;m suspecting something is going &#8220;wrong&#8221; in the training process.  The
model is greatly rewarded for succeeding on long-horizon tasks, but presumably there
is very little punishing going on for &#8220;shitty code.&#8221;  The apparent result is
that Astra is amazing <a href="https://developers.openai.com/blog/how-to-build-games-with-astra">at producing 3D stuff
</a> and it
can keep going for a very long time, coming up with its own work in the process.
I had it do quite a bit of reverse engineering of my robot vacuum in ways that
were quite impressive.  So it&#8217;s definitely cool!</p>
<h2>Codegolf Tool Calls</h2>
<p>The first issue I have with Astra comes from the type of code that it uses for
tool calls.  Codex increasingly has been relying on &#8220;just bash&#8221; to do more and
more operations.  For a few versions now the original Codex harness just uses
<code>sed</code> and other tools to read files.  You just usually can&#8217;t see them because Codex
<a href="https://github.com/openai/codex/blob/f1aac1e885f676a1129f2da0c46a3dba86392fc6/codex-rs/shell-command/src/parse_command.rs#L2290-L2504">parses the bash
commands</a>
and hides them if it recognizes them.  But Astra … really loves Python?  That is
not much of a surprise because even older OpenAI models had a tendency to
sometimes use on-demand Python code to read and manipulate files at times, but
Astra does it really quite excessively for me.</p>
<p>Now here is an important disclaimer: this project is <em>very meta</em> here because I
worked
<em>on</em> the CPython interpreter.  But I can assure you that I have seen this model
do weird Python things even in TypeScript code in Pi.  But I have the most
evidence of odd code from when I had the thing work over the weekend
with <em>zero</em> oversight from my slop factory.</p>
<p>That it writes Python is not interesting; the type of Python is interesting, and
I collected some outputs for you to gloss over.</p>
<details><summary>Python string splicing to edit C code</summary>
<p>In the Codex harness I found multiple cases where subagents resorted fully to
manual string manipulation with Python instead of using the patch tool.</p>
<div class="highlight"><pre><span></span><span class="n">python3</span> <span class="o">-</span> <span class="o">&lt;&lt;</span><span class="s1">&#39;PY&#39;</span>
<span class="kn">from</span><span class="w"> </span><span class="nn">pathlib</span><span class="w"> </span><span class="kn">import</span> <span class="n">Path</span>
<span class="n">p</span><span class="o">=</span><span class="n">Path</span><span class="p">(</span><span class="s1">&#39;Include/internal/pycore_intrinsics.h&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">p</span><span class="o">.</span><span class="n">read_text</span><span class="p">()</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="s1">&#39;#define MAX_INTRINSIC_1                         14&#39;</span><span class="p">,</span><span class="s1">&#39;#define INTRINSIC_RETAIN_ANNOTATION_CELLS        15</span><span class="se">\n\n</span><span class="s1">#define MAX_INTRINSIC_1                         15&#39;</span><span class="p">);</span><span class="n">p</span><span class="o">.</span><span class="n">write_text</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
<span class="n">p</span><span class="o">=</span><span class="n">Path</span><span class="p">(</span><span class="s1">&#39;Python/intrinsics.c&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">p</span><span class="o">.</span><span class="n">read_text</span><span class="p">();</span><span class="n">idx</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">index</span><span class="p">(</span><span class="s1">&#39;#define INTRINSIC_FUNC_ENTRY&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="p">[:</span><span class="n">idx</span><span class="p">]</span><span class="o">+</span><span class="s1">&#39;&#39;&#39;/* Hold every old cell until the compiler has published the entire site&#39;s new</span>
<span class="s1">   capture. A replaced cell&#39;s finalizer may reenter module __annotate__. */</span>
<span class="s1">static PyObject *</span>
<span class="s1">retain_annotation_cells(PyThreadState *tstate, PyObject *holders)</span>
<span class="s1">{</span>
<span class="s1">    if (!PyTuple_CheckExact(holders)) {</span>
<span class="s1">        PyErr_SetString(PyExc_TypeError, &quot;annotation holders must be a tuple&quot;);</span>
<span class="s1">        return NULL;</span>
<span class="s1">    }</span>
<span class="s1">    Py_ssize_t size = PyTuple_GET_SIZE(holders);</span>
<span class="s1">    PyObject *previous = PyTuple_New(size);</span>
<span class="s1">    if (previous == NULL) return NULL;</span>
<span class="s1">    for (Py_ssize_t i = 0; i &lt; size; i++) {</span>
<span class="s1">        PyObject *holder = PyTuple_GET_ITEM(holders, i);</span>
<span class="s1">        if (!PyCell_Check(holder)) {</span>
<span class="s1">            Py_DECREF(previous);</span>
<span class="s1">            PyErr_SetString(PyExc_TypeError, &quot;annotation holder must be a cell&quot;);</span>
<span class="s1">            return NULL;</span>
<span class="s1">        }</span>
<span class="s1">        PyObject *cell = PyCell_Get(holder);</span>
<span class="s1">        PyTuple_SET_ITEM(previous, i, cell == NULL ? Py_NewRef(Py_None) : cell);</span>
<span class="s1">    }</span>
<span class="s1">    return previous;</span>
<span class="s1">}</span>

<span class="s1">&#39;&#39;&#39;</span> <span class="o">+</span><span class="n">s</span><span class="p">[</span><span class="n">idx</span><span class="p">:];</span><span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="s1">&#39;    INTRINSIC_FUNC_ENTRY(INTRINSIC_AWAIT_BLOCK, await_block)&#39;</span><span class="p">,</span><span class="s1">&#39;    INTRINSIC_FUNC_ENTRY(INTRINSIC_AWAIT_BLOCK, await_block)</span><span class="se">\n</span><span class="s1">    INTRINSIC_FUNC_ENTRY(INTRINSIC_RETAIN_ANNOTATION_CELLS, retain_annotation_cells)&#39;</span><span class="p">);</span><span class="n">p</span><span class="o">.</span><span class="n">write_text</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
<span class="n">p</span><span class="o">=</span><span class="n">Path</span><span class="p">(</span><span class="s1">&#39;Python/codegen.c&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">p</span><span class="o">.</span><span class="n">read_text</span><span class="p">();</span><span class="n">idx</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">index</span><span class="p">(</span><span class="s1">&#39;static int</span><span class="se">\n</span><span class="s1">codegen_annassign(&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="p">[:</span><span class="n">idx</span><span class="p">]</span><span class="o">+</span><span class="s1">&#39;&#39;&#39;static int</span>
<span class="s1">codegen_retain_annotation_cells(compiler *c, location loc, PyObject *captures)</span>
<span class="s1">{</span>
<span class="s1">    Py_ssize_t pos = 0;</span>
<span class="s1">    PyObject *binding, *holder;</span>
<span class="s1">    while (PyDict_Next(captures, &amp;pos, &amp;binding, &amp;holder)) {</span>
<span class="s1">        ADDOP_NAME(c, loc, LOAD_CLOSURE, holder, cellvars);</span>
<span class="s1">    }</span>
<span class="s1">    ADDOP_I(c, loc, BUILD_TUPLE, PyDict_GET_SIZE(captures));</span>
<span class="s1">    ADDOP_I(c, loc, CALL_INTRINSIC_1, INTRINSIC_RETAIN_ANNOTATION_CELLS);</span>
<span class="s1">    return SUCCESS;</span>
<span class="s1">}</span>

<span class="s1">&#39;&#39;&#39;</span> <span class="o">+</span><span class="n">s</span><span class="p">[</span><span class="n">idx</span><span class="p">:]</span>
<span class="n">a</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">index</span><span class="p">(</span><span class="s1">&#39;                if (conditional_annotation_index != NULL) {&#39;</span><span class="p">,</span><span class="n">s</span><span class="o">.</span><span class="n">index</span><span class="p">(</span><span class="s1">&#39;codegen_annassign(compiler *c&#39;</span><span class="p">))</span>
<span class="n">b</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">index</span><span class="p">(</span><span class="s1">&#39;                if (captures != NULL) {&#39;</span><span class="p">,</span><span class="n">a</span><span class="p">)</span>
<span class="c1"># Move lookup before conditional registration and retain old cells before anything changes.</span>
<span class="n">lookupstart</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">index</span><span class="p">(</span><span class="s1">&#39;                PyObject *captures = _PyCompile_AnnotationCaptures&#39;</span><span class="p">,</span><span class="n">a</span><span class="p">)</span>
<span class="n">lookup</span><span class="o">=</span><span class="n">s</span><span class="p">[</span><span class="n">lookupstart</span><span class="p">:</span><span class="n">b</span><span class="p">]</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="s1">&#39;                    return ERROR;&#39;</span><span class="p">,</span><span class="s1">&#39;                    Py_XDECREF(conditional_annotation_index); return ERROR;&#39;</span><span class="p">)</span>
<span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="p">[:</span><span class="n">lookupstart</span><span class="p">]</span><span class="o">+</span><span class="n">s</span><span class="p">[</span><span class="n">b</span><span class="p">:]</span>
<span class="n">setup</span><span class="o">=</span><span class="n">lookup</span><span class="o">+</span><span class="s1">&#39;&#39;&#39;                if (captures != NULL &amp;&amp; codegen_retain_annotation_cells(c, loc, captures) &lt; 0) {</span>
<span class="s1">                    Py_XDECREF(conditional_annotation_index); return ERROR;</span>
<span class="s1">                }</span>
<span class="s1">&#39;&#39;&#39;</span>
<span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="p">[:</span><span class="n">a</span><span class="p">]</span><span class="o">+</span><span class="n">setup</span><span class="o">+</span><span class="n">s</span><span class="p">[</span><span class="n">a</span><span class="p">:]</span>
<span class="n">needle</span><span class="o">=</span><span class="s1">&#39;                        ADDOP_NAME(c, loc, STORE_DEREF, holder, cellvars);</span><span class="se">\n</span><span class="s1">                    }</span><span class="se">\n</span><span class="s1">                }&#39;</span>
<span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="n">needle</span><span class="p">,</span><span class="s1">&#39;                        ADDOP_NAME(c, loc, STORE_DEREF, holder, cellvars);</span><span class="se">\n</span><span class="s1">                    }</span><span class="se">\n</span><span class="s1">                    ADDOP(c, loc, POP_TOP); /* release old cells after full publication */</span><span class="se">\n</span><span class="s1">                }&#39;</span><span class="p">,</span><span class="mi">1</span><span class="p">);</span><span class="n">p</span><span class="o">.</span><span class="n">write_text</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
<span class="n">p</span><span class="o">=</span><span class="n">Path</span><span class="p">(</span><span class="s1">&#39;Include/internal/pycore_magic_number.h&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">p</span><span class="o">.</span><span class="n">read_text</span><span class="p">()</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="s1">&#39;    Python 3.16a1 3709 (Checked deferred annotation closure capture)&#39;</span><span class="p">,</span><span class="s1">&#39;    Python 3.16a1 3709 (Checked deferred annotation closure capture)</span><span class="se">\n</span><span class="s1">    Python 3.16a1 3710 (Retain replaced annotation captures until publication)&#39;</span><span class="p">)</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="s1">&#39;#define PYC_MAGIC_NUMBER 3709&#39;</span><span class="p">,</span><span class="s1">&#39;#define PYC_MAGIC_NUMBER 3710&#39;</span><span class="p">);</span><span class="n">p</span><span class="o">.</span><span class="n">write_text</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
<span class="n">p</span><span class="o">=</span><span class="n">Path</span><span class="p">(</span><span class="s1">&#39;Lib/test/test_block_annotation_captures.py&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">p</span><span class="o">.</span><span class="n">read_text</span><span class="p">();</span><span class="n">idx</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">index</span><span class="p">(</span><span class="s1">&#39;    def test_typing_consumers&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="p">[:</span><span class="n">idx</span><span class="p">]</span><span class="o">+</span><span class="s1">&#39;&#39;&#39;    def test_replaced_cell_finalizer_sees_complete_site_publication(self):</span>
<span class="s1">        module=execute(&quot;&quot;&quot;</span><span class="se">\\</span>
<span class="s1">            events=[]</span>
<span class="s1">            class V:</span>
<span class="s1">                def __init__(self,n): self.n=n</span>
<span class="s1">                def __del__(self):</span>
<span class="s1">                    if self.n == 0: events.append(__annotate__(1))</span>
<span class="s1">            for i in range(2):</span>
<span class="s1">                x=V(i) # bind x y</span>
<span class="s1">                y=i</span>
<span class="s1">                value:(x.n,y)</span>
<span class="s1">        &quot;&quot;&quot;)</span>
<span class="s1">        self.assertEqual(module.events,[{&#39;value&#39;:(1,1)}])</span>
<span class="s1">        self.assertEqual(module.__annotate__(1),{&#39;value&#39;:(1,1)})</span>

<span class="s1">&#39;&#39;&#39;</span> <span class="o">+</span><span class="n">s</span><span class="p">[</span><span class="n">idx</span><span class="p">:];</span><span class="n">p</span><span class="o">.</span><span class="n">write_text</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
<span class="n">PY</span>
<span class="n">make</span> <span class="o">-</span><span class="n">j1</span> <span class="o">&gt;</span> <span class="o">/</span><span class="n">tmp</span><span class="o">/</span><span class="n">block</span><span class="o">-</span><span class="n">annotations</span><span class="o">-</span><span class="n">build7</span><span class="o">.</span><span class="n">log</span> <span class="mi">2</span><span class="o">&gt;&amp;</span><span class="mi">1</span>
</pre></div>
</details>
<details><summary>Socket codegolf</summary>
<p>In the middle of a conversation the agent ran into &#8220;Bad file descriptor&#8221; on a test and Astra decided it needs to see if file descriptors can be passed over Unix sockets on macOS in a super compressed manner:</p>
<div class="highlight"><pre><span></span><span class="o">/</span><span class="n">usr</span><span class="o">/</span><span class="nb">bin</span><span class="o">/</span><span class="n">python3</span> <span class="o">-</span> <span class="o">&lt;&lt;</span><span class="s1">&#39;PY&#39;</span>
<span class="kn">import</span><span class="w"> </span><span class="nn">socket</span><span class="o">,</span><span class="nn">os</span><span class="o">,</span><span class="nn">array</span>
<span class="k">for</span> <span class="n">into</span> <span class="ow">in</span> <span class="p">(</span><span class="kc">False</span><span class="p">,</span><span class="kc">True</span><span class="p">):</span>
 <span class="n">a</span><span class="p">,</span><span class="n">b</span><span class="o">=</span><span class="n">socket</span><span class="o">.</span><span class="n">socketpair</span><span class="p">();</span><span class="n">fd</span><span class="o">=</span><span class="n">os</span><span class="o">.</span><span class="n">open</span><span class="p">(</span><span class="n">os</span><span class="o">.</span><span class="n">devnull</span><span class="p">,</span><span class="n">os</span><span class="o">.</span><span class="n">O_RDONLY</span><span class="p">);</span><span class="n">b</span><span class="o">.</span><span class="n">sendmsg</span><span class="p">([</span><span class="sa">b</span><span class="s1">&#39;c&#39;</span><span class="p">],[(</span><span class="n">socket</span><span class="o">.</span><span class="n">SOL_SOCKET</span><span class="p">,</span><span class="n">socket</span><span class="o">.</span><span class="n">SCM_RIGHTS</span><span class="p">,</span><span class="n">array</span><span class="o">.</span><span class="n">array</span><span class="p">(</span><span class="s1">&#39;i&#39;</span><span class="p">,[</span><span class="n">fd</span><span class="p">]))]);</span><span class="nb">print</span><span class="p">(</span><span class="s1">&#39;fds&#39;</span><span class="p">,</span><span class="n">a</span><span class="o">.</span><span class="n">fileno</span><span class="p">(),</span><span class="n">b</span><span class="o">.</span><span class="n">fileno</span><span class="p">(),</span><span class="n">fd</span><span class="p">)</span>
 <span class="k">if</span> <span class="n">into</span><span class="p">:</span><span class="n">r</span><span class="o">=</span><span class="n">a</span><span class="o">.</span><span class="n">recvmsg_into</span><span class="p">([</span><span class="nb">bytearray</span><span class="p">(</span><span class="mi">1</span><span class="p">),</span><span class="nb">bytearray</span><span class="p">(),</span><span class="nb">bytearray</span><span class="p">(</span><span class="mi">19</span><span class="p">)],</span><span class="n">socket</span><span class="o">.</span><span class="n">CMSG_SPACE</span><span class="p">(</span><span class="mi">4</span><span class="p">),</span><span class="n">socket</span><span class="o">.</span><span class="n">MSG_PEEK</span><span class="o">|</span><span class="n">socket</span><span class="o">.</span><span class="n">MSG_DONTWAIT</span><span class="p">)</span>
 <span class="k">else</span><span class="p">:</span><span class="n">r</span><span class="o">=</span><span class="n">a</span><span class="o">.</span><span class="n">recvmsg</span><span class="p">(</span><span class="mi">20</span><span class="p">,</span><span class="n">socket</span><span class="o">.</span><span class="n">CMSG_SPACE</span><span class="p">(</span><span class="mi">4</span><span class="p">),</span><span class="n">socket</span><span class="o">.</span><span class="n">MSG_PEEK</span><span class="o">|</span><span class="n">socket</span><span class="o">.</span><span class="n">MSG_DONTWAIT</span><span class="p">)</span>
 <span class="nb">print</span><span class="p">(</span><span class="s1">&#39;peek&#39;</span><span class="p">,</span><span class="n">r</span><span class="p">,</span><span class="n">flush</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span>
 <span class="n">rights</span><span class="o">=</span><span class="n">array</span><span class="o">.</span><span class="n">array</span><span class="p">(</span><span class="s1">&#39;i&#39;</span><span class="p">,</span><span class="n">r</span><span class="p">[</span><span class="mi">1</span><span class="p">][</span><span class="mi">0</span><span class="p">][</span><span class="mi">2</span><span class="p">]);</span><span class="nb">print</span><span class="p">(</span><span class="s1">&#39;rights&#39;</span><span class="p">,</span><span class="n">rights</span><span class="p">,</span><span class="n">flush</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span>
 <span class="k">for</span> <span class="n">f</span> <span class="ow">in</span> <span class="n">rights</span><span class="p">:</span>
  <span class="k">try</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="s1">&#39;stat&#39;</span><span class="p">,</span><span class="n">os</span><span class="o">.</span><span class="n">fstat</span><span class="p">(</span><span class="n">f</span><span class="p">))</span>
  <span class="k">except</span> <span class="ne">Exception</span> <span class="k">as</span> <span class="n">e</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="s1">&#39;error&#39;</span><span class="p">,</span><span class="n">e</span><span class="p">)</span>
 <span class="n">r</span><span class="o">=</span><span class="n">a</span><span class="o">.</span><span class="n">recvmsg</span><span class="p">(</span><span class="mi">20</span><span class="p">,</span><span class="n">socket</span><span class="o">.</span><span class="n">CMSG_SPACE</span><span class="p">(</span><span class="mi">4</span><span class="p">),</span><span class="n">socket</span><span class="o">.</span><span class="n">MSG_DONTWAIT</span><span class="p">);</span><span class="nb">print</span><span class="p">(</span><span class="s1">&#39;consumed&#39;</span><span class="p">,</span><span class="n">r</span><span class="p">,</span><span class="n">flush</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span>
 <span class="n">a</span><span class="o">.</span><span class="n">close</span><span class="p">();</span><span class="n">b</span><span class="o">.</span><span class="n">close</span><span class="p">();</span><span class="n">os</span><span class="o">.</span><span class="n">close</span><span class="p">(</span><span class="n">fd</span><span class="p">)</span>
<span class="n">PY</span>
</pre></div>
</details>
<details><summary>Python for agent notes patching</summary>
<p>The agent notes were rather consistently updated with Python:</p>
<div class="highlight"><pre><span></span><span class="n">python3</span> <span class="o">-</span> <span class="o">&lt;&lt;</span><span class="s1">&#39;PY&#39;</span>
<span class="kn">from</span><span class="w"> </span><span class="nn">pathlib</span><span class="w"> </span><span class="kn">import</span> <span class="n">Path</span>
<span class="n">p</span><span class="o">=</span><span class="n">Path</span><span class="p">(</span><span class="s1">&#39;agent-notes/live/block-with-bindings.md&#39;</span><span class="p">);</span><span class="n">s</span><span class="o">=</span><span class="n">p</span><span class="o">.</span><span class="n">read_text</span><span class="p">()</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="s1">&#39;  has117/77/zero failures.&#39;</span><span class="p">,</span> <span class="s1">&#39;  has117/77/zero failures; existing bundled Expat unreachable-fallthrough</span><span class="se">\n</span><span class="s1">  warnings are the only optimized warnings.&#39;</span><span class="p">)</span>
<span class="c1"># Keep the final evidence readable without rewriting historical parent requirements.</span>
<span class="k">for</span> <span class="n">a</span><span class="p">,</span><span class="n">b</span> <span class="ow">in</span> <span class="p">[(</span><span class="s1">&#39;all328&#39;</span><span class="p">,</span><span class="s1">&#39;all 328&#39;</span><span class="p">),(</span><span class="s1">&#39;pass31&#39;</span><span class="p">,</span><span class="s1">&#39;pass 31&#39;</span><span class="p">),(</span><span class="s1">&#39;pass all328&#39;</span><span class="p">,</span><span class="s1">&#39;pass all 328&#39;</span><span class="p">),(</span><span class="s1">&#39;pass,9.2s&#39;</span><span class="p">,</span><span class="s1">&#39;pass, 9.2s&#39;</span><span class="p">),(</span><span class="s1">&#39;log`,210&#39;</span><span class="p">,</span><span class="s1">&#39;log`, 210&#39;</span><span class="p">),(</span><span class="s1">&#39;log`,5,731&#39;</span><span class="p">,</span><span class="s1">&#39;log`, 5,731&#39;</span><span class="p">),(</span><span class="s1">&#39;log`:18/18&#39;</span><span class="p">,</span><span class="s1">&#39;log`: 18/18&#39;</span><span class="p">),(</span><span class="s1">&#39;pass,88&#39;</span><span class="p">,</span><span class="s1">&#39;pass, 88&#39;</span><span class="p">),(</span><span class="s1">&#39;pass,90&#39;</span><span class="p">,</span><span class="s1">&#39;pass, 90&#39;</span><span class="p">),(</span><span class="s1">&#39;skips,1m&#39;</span><span class="p">,</span><span class="s1">&#39;skips, 1m&#39;</span><span class="p">),(</span><span class="s1">&#39;all6,280&#39;</span><span class="p">,</span><span class="s1">&#39;all 6,280&#39;</span><span class="p">),(</span><span class="s1">&#39;has117&#39;</span><span class="p">,</span><span class="s1">&#39;has 117&#39;</span><span class="p">)]:</span> <span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="n">a</span><span class="p">,</span><span class="n">b</span><span class="p">)</span>
<span class="n">s</span> <span class="o">+=</span> <span class="s1">&#39;</span><span class="se">\n</span><span class="s1">Key source review: Python/symtable.c:603 (discovery), :3985 (sequential header traversal),</span><span class="se">\n</span><span class="s1">Python/codegen.c:3488 (source-only exclusion), :5836 (publication), :5853 (normal/</span><span class="se">\n</span><span class="s1">unwind reference cleanup), :5925/:6037 (enter-protected target setup).</span><span class="se">\n</span><span class="s1">&#39;</span>
<span class="n">p</span><span class="o">.</span><span class="n">write_text</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
<span class="k">for</span> <span class="n">name</span> <span class="ow">in</span> <span class="p">(</span><span class="s1">&#39;STATE.md&#39;</span><span class="p">,</span><span class="s1">&#39;build-and-test.md&#39;</span><span class="p">):</span>
 <span class="n">p</span><span class="o">=</span><span class="n">Path</span><span class="p">(</span><span class="s1">&#39;agent-notes/live&#39;</span><span class="p">)</span><span class="o">/</span><span class="n">name</span><span class="p">;</span><span class="n">s</span><span class="o">=</span><span class="n">p</span><span class="o">.</span><span class="n">read_text</span><span class="p">()</span>
 <span class="k">for</span> <span class="n">a</span><span class="p">,</span><span class="n">b</span> <span class="ow">in</span> <span class="p">[(</span><span class="s1">&#39;build:117&#39;</span><span class="p">,</span><span class="s1">&#39;build: 117&#39;</span><span class="p">),(</span><span class="s1">&#39;paths.18&#39;</span><span class="p">,</span><span class="s1">&#39;paths. 18&#39;</span><span class="p">),(</span><span class="s1">&#39;paths.</span><span class="se">\n</span><span class="s1">18&#39;</span><span class="p">,</span><span class="s1">&#39;paths.</span><span class="se">\n</span><span class="s1">18&#39;</span><span class="p">),(</span><span class="s1">&#39;and210&#39;</span><span class="p">,</span><span class="s1">&#39;and 210&#39;</span><span class="p">),(</span><span class="s1">&#39;pass5,731&#39;</span><span class="p">,</span><span class="s1">&#39;pass 5,731&#39;</span><span class="p">),(</span><span class="s1">&#39;All6,280&#39;</span><span class="p">,</span><span class="s1">&#39;All 6,280&#39;</span><span class="p">),(</span><span class="s1">&#39;failures,31&#39;</span><span class="p">,</span><span class="s1">&#39;failures, 31&#39;</span><span class="p">),(</span><span class="s1">&#39;in</span><span class="se">\n</span><span class="s1">115s&#39;</span><span class="p">,</span><span class="s1">&#39;in</span><span class="se">\n</span><span class="s1">115s&#39;</span><span class="p">),(</span><span class="s1">&#39;have117&#39;</span><span class="p">,</span><span class="s1">&#39;have 117&#39;</span><span class="p">),(</span><span class="s1">&#39;paths.</span><span class="se">\n</span><span class="s1">18&#39;</span><span class="p">,</span><span class="s1">&#39;paths.</span><span class="se">\n</span><span class="s1">18&#39;</span><span class="p">),(</span><span class="s1">&#39;18 focused,210&#39;</span><span class="p">,</span><span class="s1">&#39;18 focused, 210&#39;</span><span class="p">),(</span><span class="s1">&#39;and5,731&#39;</span><span class="p">,</span><span class="s1">&#39;and 5,731&#39;</span><span class="p">),(</span><span class="s1">&#39;all6,280&#39;</span><span class="p">,</span><span class="s1">&#39;all 6,280&#39;</span><span class="p">)]:</span> <span class="n">s</span><span class="o">=</span><span class="n">s</span><span class="o">.</span><span class="n">replace</span><span class="p">(</span><span class="n">a</span><span class="p">,</span><span class="n">b</span><span class="p">)</span>
 <span class="n">p</span><span class="o">.</span><span class="n">write_text</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
<span class="n">PY</span>
<span class="n">git</span> <span class="n">diff</span> <span class="o">--</span><span class="n">check</span>
<span class="n">git</span> <span class="n">add</span> <span class="o">-</span><span class="n">u</span>
<span class="n">git</span> <span class="n">add</span> <span class="n">Lib</span><span class="o">/</span><span class="n">test</span><span class="o">/</span><span class="n">test_block_with_bindings</span><span class="o">.</span><span class="n">py</span> <span class="n">agent</span><span class="o">-</span><span class="n">notes</span><span class="o">/</span><span class="n">done</span><span class="o">/</span><span class="n">asyncio</span><span class="o">-</span><span class="n">task</span><span class="o">-</span><span class="n">drivers</span><span class="o">.</span><span class="n">md</span>
<span class="n">git</span> <span class="n">diff</span> <span class="o">--</span><span class="n">cached</span> <span class="o">--</span><span class="n">stat</span>
<span class="n">git</span> <span class="n">commit</span> <span class="o">-</span><span class="n">m</span> <span class="s1">&#39;Add explicit with and async with header bindings&#39;</span>
</pre></div>
</details>
<details><summary>Using Python to run Node.js</summary>
<p>In multiple cases it used Python to spawn Node.js on another machine.  It first wrote the script, then it used Bash to run Python, then that program ran Node.js via <code>prlctl</code> on my Windows box.</p>
<div class="highlight"><pre><span></span><span class="kn">import</span><span class="w"> </span><span class="nn">subprocess</span>
<span class="n">code</span> <span class="o">=</span> <span class="s2">&quot;const</span><span class="si">{readFileSync}</span><span class="s2">=require(&#39;fs&#39;);const{strict:a}=require(&#39;assert&#39;);const c=require(&#39;C:/Users/mitsuhiko/AppData/Local/Temp/pi-clipboard-threads/win32-arm64.node&#39;);(async()=&gt;{const p=c.getText();a.ok(p instanceof Promise);const saved=await p;const image=await c.getImage();if(image||saved===null){console.log(&#39;arm64 async text/image reads passed; preserving non-text clipboard&#39;);return}try{for(const text of [&#39;café 日本語&#39;,&#39;&#39;, &#39;large&#39;.repeat(200000)]){const p=c.setText(text);a.ok(p instanceof Promise);await p;a.equal(await c.getText(),text);a.equal(await c.getImage(),null)}console.log(&#39;Windows ARM64 async Unicode, empty, large text and empty image passed&#39;)}finally{await c.setText(saved)}})().catch(e=&gt;{console.error(e);process.exitCode=1})&quot;</span>
<span class="n">subprocess</span><span class="o">.</span><span class="n">run</span><span class="p">([</span><span class="s1">&#39;prlctl&#39;</span><span class="p">,</span> <span class="s1">&#39;exec&#39;</span><span class="p">,</span> <span class="s1">&#39;Windows 11&#39;</span><span class="p">,</span> <span class="s1">&#39;--current-user&#39;</span><span class="p">,</span> <span class="s1">&#39;C:</span><span class="se">\\</span><span class="s1">Program Files</span><span class="se">\\</span><span class="s1">nodejs</span><span class="se">\\</span><span class="s1">node.exe&#39;</span><span class="p">,</span> <span class="s1">&#39;-e&#39;</span><span class="p">,</span> <span class="n">code</span><span class="p">],</span> <span class="n">check</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span>
</pre></div>
</details>
<details><summary>Python to run Node.js to run PowerShell</summary>
<p>Since it was already doing that, it used Bash to run Python to then run Node.js to then use Node.js to invoke PowerShell.</p>
<div class="highlight"><pre><span></span><span class="kn">import</span><span class="w"> </span><span class="nn">subprocess</span>
<span class="n">code</span> <span class="o">=</span> <span class="s2">&quot;process.env.PSModulePath=&#39;C:/Windows/System32/WindowsPowerShell/v1.0/Modules&#39;;require(&#39;child_process&#39;).spawnSync(&#39;powershell.exe&#39;,[&#39;-NoProfile&#39;,&#39;-NonInteractive&#39;,&#39;-ExecutionPolicy&#39;,&#39;Bypass&#39;,&#39;-File&#39;,&#39;C:/Users/mitsuhiko/AppData/Local/Temp/pi-clipboard-threads/pi-clipboard-windows.ps1&#39;],{stdio:&#39;inherit&#39;});console.log(&#39;completed&#39;)&quot;</span>
<span class="n">subprocess</span><span class="o">.</span><span class="n">run</span><span class="p">([</span><span class="s1">&#39;prlctl&#39;</span><span class="p">,</span> <span class="s1">&#39;exec&#39;</span><span class="p">,</span> <span class="s1">&#39;Windows 11&#39;</span><span class="p">,</span> <span class="s1">&#39;--current-user&#39;</span><span class="p">,</span> <span class="s1">&#39;C:</span><span class="se">\\</span><span class="s1">Program Files</span><span class="se">\\</span><span class="s1">nodejs</span><span class="se">\\</span><span class="s1">node.exe&#39;</span><span class="p">,</span> <span class="s1">&#39;-e&#39;</span><span class="p">,</span> <span class="n">code</span><span class="p">],</span> <span class="n">check</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span>
</pre></div>
</details>
<p>You can consider this amusing, but I have some questions here.  The first
problem with this is that it&#8217;s unreadable for a human.  If you wanna follow
along with what is going on, then good luck.  Particularly once it opts out of
using the edit tools that the harness provides, you&#8217;re going to have to resort
to using the diff viewer of the final artifacts since it&#8217;s almost impossible to
visualize the changes as they happen by reading the code.</p>
<p>This is not quite as bad in Pi for the most part because I mostly see it editing
with the <code>edit</code> tool.  When however goes all bananza with subagents (where the
agent believes nobody is looking) it&#8217;s resorting to all kinds of increasingly
bizarre behavior.  I actually don&#8217;t know if the model thinks someone is looking,
but that&#8217;s the vibe I&#8217;m getting.</p>
<p>But then it starts doing the same nonsense in code that actually gets committed.
I have mostly seen this in tests, but you can also see this for instance when
it writes JavaScript or CSS embedded in HTML.  It almost seems like when it&#8217;s
&#8220;one step removed&#8221; from regular code, it starts falling into these patterns.</p>
<p>Here are some unit tests that it created:</p>
<details><summary>Complete disregard for whitespace and indentation</summary>
<div class="highlight"><pre><span></span><span class="k">def</span><span class="w"> </span><span class="nf">test_unpack_suspension_and_continuation_close</span><span class="p">(</span><span class="bp">self</span><span class="p">):</span>
    <span class="kn">from</span><span class="w"> </span><span class="nn">continuations</span><span class="w"> </span><span class="kn">import</span> <span class="n">Continuation</span><span class="p">,</span><span class="n">suspend</span>
    <span class="n">readers</span><span class="o">=</span><span class="p">[]</span>
    <span class="k">class</span><span class="w"> </span><span class="nc">Source</span><span class="p">:</span>
        <span class="k">def</span><span class="w"> </span><span class="fm">__iter__</span><span class="p">(</span><span class="bp">self</span><span class="p">):</span>
            <span class="k">yield</span> <span class="mi">1</span>
            <span class="n">suspend</span><span class="p">(</span><span class="s1">&#39;unpacking&#39;</span><span class="p">)</span>
            <span class="k">yield</span> <span class="mi">2</span>
    <span class="n">ns</span><span class="o">=</span><span class="n">execute</span><span class="p">(</span><span class="s1">&#39;&#39;&#39;</span>
<span class="s1">        def run():</span>
<span class="s1">            a,b=&#39;old-a&#39;,&#39;old-b&#39;</span>
<span class="s1">            readers.append(lambda: (a,b))</span>
<span class="s1">            def a,b=Source()</span>
<span class="s1">            suspend(&#39;published&#39;)</span>
<span class="s1">    &#39;&#39;&#39;</span><span class="p">,</span><span class="n">Source</span><span class="o">=</span><span class="n">Source</span><span class="p">,</span><span class="n">readers</span><span class="o">=</span><span class="n">readers</span><span class="p">,</span><span class="n">suspend</span><span class="o">=</span><span class="n">suspend</span><span class="p">)</span>
    <span class="k">with</span> <span class="n">Continuation</span><span class="p">(</span><span class="n">ns</span><span class="p">[</span><span class="s1">&#39;run&#39;</span><span class="p">])</span> <span class="k">as</span> <span class="n">continuation</span><span class="p">:</span>
        <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">continuation</span><span class="o">.</span><span class="n">resume</span><span class="p">(),</span><span class="s1">&#39;unpacking&#39;</span><span class="p">)</span>
        <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">readers</span><span class="p">[</span><span class="mi">0</span><span class="p">](),(</span><span class="s1">&#39;old-a&#39;</span><span class="p">,</span><span class="s1">&#39;old-b&#39;</span><span class="p">))</span>
        <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">continuation</span><span class="o">.</span><span class="n">resume</span><span class="p">(),</span><span class="s1">&#39;published&#39;</span><span class="p">)</span>
        <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">readers</span><span class="p">[</span><span class="mi">0</span><span class="p">](),(</span><span class="mi">1</span><span class="p">,</span><span class="mi">2</span><span class="p">))</span>
    <span class="k">class</span><span class="w"> </span><span class="nc">Value</span><span class="p">:</span><span class="k">pass</span>
    <span class="n">refs</span><span class="o">=</span><span class="p">[];</span><span class="n">frames</span><span class="o">=</span><span class="p">[];</span><span class="n">callbacks</span><span class="o">=</span><span class="p">[]</span>
    <span class="n">ns</span><span class="o">=</span><span class="n">execute</span><span class="p">(</span><span class="s1">&#39;&#39;&#39;</span>
<span class="s1">        def run():</span>
<span class="s1">            for def x in [Value()]:</span>
<span class="s1">                refs.append(weakref.ref(x))</span>
<span class="s1">                frames.append(sys._getframe())</span>
<span class="s1">                callbacks.append(lambda: x)</span>
<span class="s1">                suspend(&#39;body&#39;)</span>
<span class="s1">    &#39;&#39;&#39;</span><span class="p">,</span><span class="n">Value</span><span class="o">=</span><span class="n">Value</span><span class="p">,</span><span class="n">refs</span><span class="o">=</span><span class="n">refs</span><span class="p">,</span><span class="n">frames</span><span class="o">=</span><span class="n">frames</span><span class="p">,</span><span class="n">callbacks</span><span class="o">=</span><span class="n">callbacks</span><span class="p">,</span><span class="n">weakref</span><span class="o">=</span><span class="n">weakref</span><span class="p">,</span><span class="n">sys</span><span class="o">=</span><span class="n">sys</span><span class="p">,</span><span class="n">suspend</span><span class="o">=</span><span class="n">suspend</span><span class="p">)</span>
    <span class="k">with</span> <span class="n">Continuation</span><span class="p">(</span><span class="n">ns</span><span class="p">[</span><span class="s1">&#39;run&#39;</span><span class="p">])</span> <span class="k">as</span> <span class="n">continuation</span><span class="p">:</span><span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">continuation</span><span class="o">.</span><span class="n">resume</span><span class="p">(),</span><span class="s1">&#39;body&#39;</span><span class="p">)</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertNotIn</span><span class="p">(</span><span class="s1">&#39;x&#39;</span><span class="p">,</span><span class="n">frames</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span><span class="o">.</span><span class="n">f_locals</span><span class="p">)</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertIsNotNone</span><span class="p">(</span><span class="n">refs</span><span class="p">[</span><span class="mi">0</span><span class="p">]());</span><span class="n">callbacks</span><span class="o">.</span><span class="n">clear</span><span class="p">();</span><span class="bp">self</span><span class="o">.</span><span class="n">assertIsNone</span><span class="p">(</span><span class="n">refs</span><span class="p">[</span><span class="mi">0</span><span class="p">]())</span>

<span class="k">def</span><span class="w"> </span><span class="nf">test_ast_roundtrips_and_future_annotation_unparse</span><span class="p">(</span><span class="bp">self</span><span class="p">):</span>
    <span class="n">source</span><span class="o">=</span><span class="s1">&#39;callback=lambda {for def a, [b,*rest] in [(1,[2,3])] {return a,b,rest}}&#39;</span>
    <span class="n">tree</span><span class="o">=</span><span class="n">ast</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="n">source</span><span class="p">);</span><span class="n">node</span><span class="o">=</span><span class="n">tree</span><span class="o">.</span><span class="n">body</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span><span class="o">.</span><span class="n">value</span><span class="o">.</span><span class="n">body</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertIsInstance</span><span class="p">(</span><span class="n">node</span><span class="p">,</span><span class="n">ast</span><span class="o">.</span><span class="n">ForBinding</span><span class="p">)</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">node</span><span class="o">.</span><span class="n">_fields</span><span class="p">,(</span><span class="s1">&#39;target&#39;</span><span class="p">,</span><span class="s1">&#39;iter&#39;</span><span class="p">,</span><span class="s1">&#39;body&#39;</span><span class="p">,</span><span class="s1">&#39;orelse&#39;</span><span class="p">,</span><span class="s1">&#39;type_comment&#39;</span><span class="p">))</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">node</span><span class="o">.</span><span class="n">lineno</span><span class="p">,</span><span class="mi">1</span><span class="p">);</span><span class="bp">self</span><span class="o">.</span><span class="n">assertGreater</span><span class="p">(</span><span class="n">node</span><span class="o">.</span><span class="n">end_col_offset</span><span class="p">,</span><span class="n">node</span><span class="o">.</span><span class="n">col_offset</span><span class="p">)</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">ast</span><span class="o">.</span><span class="n">dump</span><span class="p">(</span><span class="n">tree</span><span class="p">),</span><span class="n">ast</span><span class="o">.</span><span class="n">dump</span><span class="p">(</span><span class="n">ast</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="n">ast</span><span class="o">.</span><span class="n">unparse</span><span class="p">(</span><span class="n">tree</span><span class="p">))))</span>
    <span class="n">ns</span><span class="o">=</span><span class="n">execute</span><span class="p">(</span><span class="s1">&#39;from __future__ import annotations</span><span class="se">\n</span><span class="s1">def f(arg: &#39;</span><span class="o">+</span><span class="n">source</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="s1">&#39;=&#39;</span><span class="p">,</span><span class="mi">1</span><span class="p">)[</span><span class="mi">1</span><span class="p">]</span><span class="o">+</span><span class="s1">&#39;): pass&#39;</span><span class="p">)</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="nb">eval</span><span class="p">(</span><span class="n">ns</span><span class="p">[</span><span class="s1">&#39;f&#39;</span><span class="p">]</span><span class="o">.</span><span class="vm">__annotations__</span><span class="p">[</span><span class="s1">&#39;arg&#39;</span><span class="p">])(),(</span><span class="mi">1</span><span class="p">,</span><span class="mi">2</span><span class="p">,[</span><span class="mi">3</span><span class="p">]))</span>
    <span class="n">tree</span><span class="o">=</span><span class="n">ast</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="s1">&#39;async def f():</span><span class="se">\n</span><span class="s1"> async for def x in values: pass # type: ignored</span><span class="se">\n</span><span class="s1">&#39;</span><span class="p">)</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertIsInstance</span><span class="p">(</span><span class="n">tree</span><span class="o">.</span><span class="n">body</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span><span class="o">.</span><span class="n">body</span><span class="p">[</span><span class="mi">0</span><span class="p">],</span><span class="n">ast</span><span class="o">.</span><span class="n">AsyncForBinding</span><span class="p">)</span>
    <span class="bp">self</span><span class="o">.</span><span class="n">assertEqual</span><span class="p">(</span><span class="n">ast</span><span class="o">.</span><span class="n">dump</span><span class="p">(</span><span class="n">tree</span><span class="p">),</span><span class="n">ast</span><span class="o">.</span><span class="n">dump</span><span class="p">(</span><span class="n">ast</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="n">ast</span><span class="o">.</span><span class="n">unparse</span><span class="p">(</span><span class="n">tree</span><span class="p">))))</span>
</pre></div>
</details>
<p>So at least in some situations, the Python slop that it normally code-golfs for
token-efficient tool calls leaks into the Python code it generates that should
be stored.  And well, it&#8217;s clearly more token efficient.  The two unit tests
above, when indented to the class structure they were in, are 10% more token
efficient in this form than after a <code>ruff format</code>.</p>
<h2>It&#8217;s AGI If You Don&#8217;t Look</h2>
<p>I think there are a handful of things happening now that are pushing the whole
thing in directions that are in conflict with one another.  The training runs
for these models are rapidly accelerating and they are now presumably also moving
towards recursive self-improvement.  The reward for the models is probably a
combination of token efficiency, task completion rate and maybe some simple
indicators like cyclomatic complexity.  But we humans don&#8217;t think of code that
is readable or understandable by simple, readily quantifiable metrics.  All
those things you can easily measure in isolation, and you can also optimize for
them quite locally.</p>
<p>But these local optimizations do not produce global optimums, and the fewer of us
are looking at the output, the less it matters.  Obviously my software factory ran
aground over the ~35 hours that it ran, but you can see the gradual regression
towards insanity from the notes that it produced.  For instance the task naming
in the task file starts with an optimistic 1, 2, 3, 5, 5a but then eventually
gets to 8a, 8a1, and then ends up with 8b2c2b3 and &#8220;8b2c2b2b checkpoint1&#8221;.  The
code that it produced got ever more wild.  I don&#8217;t want to bore you with what
it tried to build, but here are some example pieces of the interpreter changes:</p>
<details><summary>Hardcoded constants everywhere</summary>
<p>I have no idea where it got those numbers from, but at one point it started
passing random constants from one module to a C implementation.  Initially that
started out as a function that it mainly needed to do test assertions, but just
before I turned off that experiment, that function started to be relied upon by
non-test code as well.</p>
<div class="highlight"><pre><span></span><span class="k">static</span><span class="w"> </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span>
<span class="nf">native_probe_run_impl</span><span class="p">(</span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">callback</span><span class="p">,</span><span class="w"> </span><span class="kt">int</span><span class="w"> </span><span class="n">sleep</span><span class="p">,</span><span class="w"> </span><span class="kt">int</span><span class="w"> </span><span class="n">operation</span><span class="p">,</span><span class="w"> </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">other</span><span class="p">)</span>
<span class="p">{</span>
<span class="w">    </span><span class="n">pthread_mutexattr_t</span><span class="w"> </span><span class="n">attr</span><span class="p">;</span>
<span class="w">    </span><span class="n">pthread_mutex_t</span><span class="w"> </span><span class="n">mutex</span><span class="p">;</span>
<span class="w">    </span><span class="n">pthread_mutexattr_init</span><span class="p">(</span><span class="o">&amp;</span><span class="n">attr</span><span class="p">);</span>
<span class="w">    </span><span class="n">pthread_mutexattr_settype</span><span class="p">(</span><span class="o">&amp;</span><span class="n">attr</span><span class="p">,</span><span class="w"> </span><span class="n">PTHREAD_MUTEX_RECURSIVE</span><span class="p">);</span>
<span class="w">    </span><span class="n">pthread_mutex_init</span><span class="p">(</span><span class="o">&amp;</span><span class="n">mutex</span><span class="p">,</span><span class="w"> </span><span class="o">&amp;</span><span class="n">attr</span><span class="p">);</span>
<span class="w">    </span><span class="n">pthread_mutexattr_destroy</span><span class="p">(</span><span class="o">&amp;</span><span class="n">attr</span><span class="p">);</span>
<span class="w">    </span><span class="n">pthread_mutex_lock</span><span class="p">(</span><span class="o">&amp;</span><span class="n">mutex</span><span class="p">);</span>
<span class="w">    </span><span class="kt">int</span><span class="w"> </span><span class="n">previous</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">native_sentinel</span><span class="p">;</span>
<span class="w">    </span><span class="n">pthread_mutex_t</span><span class="w"> </span><span class="o">*</span><span class="n">previous_mutex</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">native_mutex</span><span class="p">;</span>
<span class="w">    </span><span class="n">native_sentinel</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">previous</span><span class="w"> </span><span class="o">+</span><span class="w"> </span><span class="mi">1</span><span class="p">;</span>
<span class="w">    </span><span class="n">native_mutex</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="o">&amp;</span><span class="n">mutex</span><span class="p">;</span>
<span class="w">    </span><span class="n">PyThreadState</span><span class="w"> </span><span class="o">*</span><span class="n">tstate</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyThreadState_Get</span><span class="p">();</span>
<span class="w">    </span><span class="n">PyGILState_STATE</span><span class="w"> </span><span class="n">gil</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyGILState_Ensure</span><span class="p">();</span>
<span class="w">    </span><span class="kt">int</span><span class="w"> </span><span class="n">saved_errno</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">errno</span><span class="p">;</span>
<span class="w">    </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nb">NULL</span><span class="p">;</span>
<span class="w">    </span><span class="n">Py_ssize_t</span><span class="w"> </span><span class="n">value</span><span class="p">;</span>
<span class="w">    </span><span class="cm">/* No intervening Python frame: these exercise ambient C provenance. */</span>
<span class="w">    </span><span class="k">switch</span><span class="w"> </span><span class="p">(</span><span class="n">operation</span><span class="p">)</span><span class="w"> </span><span class="p">{</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">0</span><span class="p">:</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyObject_CallNoArgs</span><span class="p">(</span><span class="n">callback</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">1</span><span class="p">:</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyNumber_Add</span><span class="p">(</span><span class="n">callback</span><span class="p">,</span><span class="w"> </span><span class="n">other</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">2</span><span class="p">:</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyNumber_Negative</span><span class="p">(</span><span class="n">callback</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">3</span><span class="p">:</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyObject_RichCompare</span><span class="p">(</span><span class="n">callback</span><span class="p">,</span><span class="w"> </span><span class="n">other</span><span class="p">,</span><span class="w"> </span><span class="n">Py_LT</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">4</span><span class="p">:</span>
<span class="w">            </span><span class="n">value</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyObject_IsTrue</span><span class="p">(</span><span class="n">callback</span><span class="p">);</span>
<span class="w">            </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">value</span><span class="w"> </span><span class="o">&gt;=</span><span class="w"> </span><span class="mi">0</span><span class="p">)</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyBool_FromLong</span><span class="p">(</span><span class="n">value</span><span class="p">);</span>
<span class="w">            </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">5</span><span class="p">:</span>
<span class="w">            </span><span class="n">value</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyObject_Length</span><span class="p">(</span><span class="n">callback</span><span class="p">);</span>
<span class="w">            </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">value</span><span class="w"> </span><span class="o">&gt;=</span><span class="w"> </span><span class="mi">0</span><span class="p">)</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyLong_FromSsize_t</span><span class="p">(</span><span class="n">value</span><span class="p">);</span>
<span class="w">            </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">6</span><span class="p">:</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyObject_GetIter</span><span class="p">(</span><span class="n">callback</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">7</span><span class="p">:</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyIter_Next</span><span class="p">(</span><span class="n">callback</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">8</span><span class="p">:</span><span class="w"> </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyObject_GetItem</span><span class="p">(</span><span class="n">callback</span><span class="p">,</span><span class="w"> </span><span class="n">other</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="cm">/* ... */</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">21</span><span class="p">:</span>
<span class="w">            </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyType_Type</span><span class="p">.</span><span class="n">tp_call</span><span class="p">(</span><span class="n">callback</span><span class="p">,</span><span class="w"> </span><span class="n">other</span><span class="p">,</span><span class="w"> </span><span class="nb">NULL</span><span class="p">);</span>
<span class="w">            </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">22</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">23</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">24</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">25</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">26</span><span class="p">:</span>
<span class="w">            </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">conversion_probe</span><span class="p">(</span><span class="n">operation</span><span class="p">,</span><span class="w"> </span><span class="n">callback</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">27</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">28</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">29</span><span class="p">:</span>
<span class="w">            </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">protocol_probe</span><span class="p">(</span><span class="n">operation</span><span class="p">,</span><span class="w"> </span><span class="n">callback</span><span class="p">,</span><span class="w"> </span><span class="n">other</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">30</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">31</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">32</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">33</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">34</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">35</span><span class="p">:</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">36</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">37</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">38</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">39</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">40</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">41</span><span class="p">:</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">42</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">43</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">44</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">45</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">46</span><span class="p">:</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">47</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">48</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">49</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">50</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">51</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">52</span><span class="p">:</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">53</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">54</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">55</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">56</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">57</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">58</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">59</span><span class="p">:</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">60</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">61</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">62</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">63</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">64</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">65</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">66</span><span class="p">:</span>
<span class="w">        </span><span class="k">case</span><span class="w"> </span><span class="mi">67</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">68</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">69</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">70</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">71</span><span class="p">:</span><span class="w"> </span><span class="k">case</span><span class="w"> </span><span class="mi">72</span><span class="p">:</span>
<span class="w">            </span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">collection_probe</span><span class="p">(</span><span class="n">operation</span><span class="p">,</span><span class="w"> </span><span class="n">callback</span><span class="p">,</span><span class="w"> </span><span class="n">other</span><span class="p">);</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">        </span><span class="k">default</span><span class="o">:</span><span class="w"> </span><span class="n">PyErr_SetString</span><span class="p">(</span><span class="n">PyExc_ValueError</span><span class="p">,</span><span class="w"> </span><span class="s">&quot;bad probe operation&quot;</span><span class="p">);</span>
<span class="w">    </span><span class="p">}</span>
</pre></div>
</details>
<details><summary>Multiple same-line macro invocations in C</summary>
<p>This code style does not exist in the CPython code base, yet it shows up in newly generated code.</p>
<div class="highlight"><pre><span></span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">info</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyTuple_Pack</span><span class="p">(</span><span class="mi">3</span><span class="p">,</span><span class="w"> </span><span class="n">name</span><span class="p">,</span><span class="w"> </span><span class="n">mangled</span><span class="p">,</span><span class="w"> </span><span class="n">suite</span><span class="o">-&gt;</span><span class="n">su_id</span><span class="p">);</span>
<span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">flags</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyLong_FromLong</span><span class="p">(</span><span class="n">DEF_LOCAL</span><span class="p">);</span>
<span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">key</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="nb">NULL</span><span class="w"> </span><span class="o">||</span><span class="w"> </span><span class="n">info</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="nb">NULL</span><span class="w"> </span><span class="o">||</span><span class="w"> </span><span class="n">flags</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="nb">NULL</span><span class="w"> </span><span class="o">||</span>
<span class="w">    </span><span class="n">PyDict_SetItem</span><span class="p">(</span><span class="n">suite</span><span class="o">-&gt;</span><span class="n">su_bindings</span><span class="p">,</span><span class="w"> </span><span class="n">mangled</span><span class="p">,</span><span class="w"> </span><span class="n">key</span><span class="p">)</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="w"> </span><span class="o">||</span>
<span class="w">    </span><span class="n">PyDict_SetItem</span><span class="p">(</span><span class="n">st</span><span class="o">-&gt;</span><span class="n">st_cur</span><span class="o">-&gt;</span><span class="n">ste_block_bindings</span><span class="p">,</span><span class="w"> </span><span class="n">key</span><span class="p">,</span><span class="w"> </span><span class="n">info</span><span class="p">)</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="w"> </span><span class="o">||</span>
<span class="w">    </span><span class="p">(</span><span class="n">private</span><span class="w"> </span><span class="o">&amp;&amp;</span><span class="w"> </span><span class="n">PyDict_SetItem</span><span class="p">(</span><span class="n">st</span><span class="o">-&gt;</span><span class="n">st_binding_info</span><span class="p">,</span><span class="w"> </span><span class="n">key</span><span class="p">,</span><span class="w"> </span><span class="n">info</span><span class="p">)</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="p">)</span><span class="w"> </span><span class="o">||</span>
<span class="w">    </span><span class="p">(</span><span class="n">private</span><span class="w"> </span><span class="o">&amp;&amp;</span><span class="w"> </span><span class="n">PyDict_SetItem</span><span class="p">(</span><span class="n">st</span><span class="o">-&gt;</span><span class="n">st_cur</span><span class="o">-&gt;</span><span class="n">ste_symbols</span><span class="p">,</span><span class="w"> </span><span class="n">key</span><span class="p">,</span><span class="w"> </span><span class="n">flags</span><span class="p">)</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="p">))</span><span class="w"> </span><span class="p">{</span>
<span class="w">    </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">mangled</span><span class="p">);</span><span class="w"> </span><span class="n">Py_XDECREF</span><span class="p">(</span><span class="n">key</span><span class="p">);</span><span class="w"> </span><span class="n">Py_XDECREF</span><span class="p">(</span><span class="n">info</span><span class="p">);</span><span class="w"> </span><span class="n">Py_XDECREF</span><span class="p">(</span><span class="n">flags</span><span class="p">);</span>
<span class="w">    </span><span class="k">goto</span><span class="w"> </span><span class="n">error</span><span class="p">;</span>
<span class="p">}</span>
<span class="n">Py_DECREF</span><span class="p">(</span><span class="n">mangled</span><span class="p">);</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">key</span><span class="p">);</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">info</span><span class="p">);</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">flags</span><span class="p">);</span>
</pre></div>
</details>
<details><summary>Random indexes in production code</summary>
<p>As with the numbers for the operators, it also uses random
integers in a list to stash away state.</p>
<div class="highlight"><pre><span></span><span class="k">def</span><span class="w"> </span><span class="nf">_register_task</span><span class="p">(</span><span class="n">task</span><span class="p">):</span>
<span class="w">    </span><span class="sd">&quot;&quot;&quot;Register an asyncio Task scheduled to run on an event loop.&quot;&quot;&quot;</span>
    <span class="n">_scheduled_tasks</span><span class="o">.</span><span class="n">add</span><span class="p">(</span><span class="n">task</span><span class="p">)</span>
    <span class="k">if</span> <span class="n">_task_accelerator</span> <span class="ow">is</span> <span class="ow">not</span> <span class="kc">None</span><span class="p">:</span>
        <span class="n">_task_accelerator</span><span class="p">[</span><span class="mi">6</span><span class="p">](</span><span class="n">task</span><span class="p">)</span>


<span class="k">def</span><span class="w"> </span><span class="nf">_register_eager_task</span><span class="p">(</span><span class="n">task</span><span class="p">):</span>
<span class="w">    </span><span class="sd">&quot;&quot;&quot;Register an asyncio Task about to be eagerly executed.&quot;&quot;&quot;</span>
    <span class="n">_eager_tasks</span><span class="o">.</span><span class="n">add</span><span class="p">(</span><span class="n">task</span><span class="p">)</span>
    <span class="k">if</span> <span class="n">_task_accelerator</span> <span class="ow">is</span> <span class="ow">not</span> <span class="kc">None</span><span class="p">:</span>
        <span class="n">_task_accelerator</span><span class="p">[</span><span class="mi">8</span><span class="p">](</span><span class="n">task</span><span class="p">)</span>


<span class="k">def</span><span class="w"> </span><span class="nf">_enter_task</span><span class="p">(</span><span class="n">loop</span><span class="p">,</span> <span class="n">task</span><span class="p">):</span>
    <span class="k">if</span> <span class="p">(</span><span class="n">_task_accelerator</span> <span class="ow">is</span> <span class="ow">not</span> <span class="kc">None</span> <span class="ow">and</span>
            <span class="n">_task_accelerator</span><span class="p">[</span><span class="mi">5</span><span class="p">]()</span> <span class="ow">is</span> <span class="n">loop</span> <span class="ow">and</span> <span class="n">loop</span> <span class="ow">not</span> <span class="ow">in</span> <span class="n">_current_tasks</span><span class="p">):</span>
        <span class="k">return</span> <span class="n">_task_accelerator</span><span class="p">[</span><span class="mi">1</span><span class="p">](</span><span class="n">loop</span><span class="p">,</span> <span class="n">task</span><span class="p">)</span>
    <span class="c1"># ...</span>
</pre></div>
</details>
<details><summary>Hideous tokenizer code in C</summary>
<p>This is not the codebase&#8217;s coding style, and quite frankly it should not be anyone&#8217;s coding style.  I do not understand what motivated the model to do this.</p>
<div class="highlight"><pre><span></span><span class="k">static</span><span class="w"> </span><span class="kt">int</span>
<span class="nf">apply_layout</span><span class="p">(</span><span class="n">tokenizeriterobject</span><span class="w"> </span><span class="o">*</span><span class="n">it</span><span class="p">)</span>
<span class="p">{</span>
<span class="w">    </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">source</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyBytes_FromStringAndSize</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">tok</span><span class="o">-&gt;</span><span class="n">source</span><span class="p">.</span><span class="n">bytes</span><span class="p">,</span><span class="w"> </span><span class="n">it</span><span class="o">-&gt;</span><span class="n">tok</span><span class="o">-&gt;</span><span class="n">source</span><span class="p">.</span><span class="n">len</span><span class="p">);</span>
<span class="w">    </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">source</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="nb">NULL</span><span class="p">)</span><span class="w"> </span><span class="k">return</span><span class="w"> </span><span class="mi">-1</span><span class="p">;</span>
<span class="w">    </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">events</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">_PyPegen_tokenize_layout</span><span class="p">(</span><span class="n">PyBytes_AS_STRING</span><span class="p">(</span><span class="n">source</span><span class="p">),</span><span class="w"> </span><span class="n">it</span><span class="o">-&gt;</span><span class="n">tok</span><span class="o">-&gt;</span><span class="n">filename</span><span class="p">);</span>
<span class="w">    </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">source</span><span class="p">);</span>
<span class="w">    </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">events</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="nb">NULL</span><span class="p">)</span><span class="w"> </span><span class="k">return</span><span class="w"> </span><span class="mi">-1</span><span class="p">;</span>
<span class="w">    </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">result</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyList_New</span><span class="p">(</span><span class="mi">0</span><span class="p">);</span>
<span class="w">    </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">result</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="nb">NULL</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">events</span><span class="p">);</span><span class="w"> </span><span class="k">return</span><span class="w"> </span><span class="mi">-1</span><span class="p">;</span><span class="w"> </span><span class="p">}</span>
<span class="w">    </span><span class="n">Py_ssize_t</span><span class="w"> </span><span class="n">index</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="mi">0</span><span class="p">;</span>
<span class="w">    </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">first_pos</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyTuple_GET_ITEM</span><span class="p">(</span><span class="n">PyList_GET_ITEM</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">,</span><span class="w"> </span><span class="mi">0</span><span class="p">),</span><span class="w"> </span><span class="mi">2</span><span class="p">);</span>
<span class="w">    </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">last_pos</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyTuple_GET_ITEM</span><span class="p">(</span><span class="n">PyList_GET_ITEM</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">,</span><span class="w"> </span><span class="n">PyList_GET_SIZE</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">)</span><span class="mi">-1</span><span class="p">),</span><span class="w"> </span><span class="mi">2</span><span class="p">);</span>
<span class="w">    </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">previous</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="nb">NULL</span><span class="p">;</span>
<span class="w">    </span><span class="k">for</span><span class="w"> </span><span class="p">(</span><span class="n">Py_ssize_t</span><span class="w"> </span><span class="n">i</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="mi">0</span><span class="p">;</span><span class="w"> </span><span class="n">i</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="n">PyList_GET_SIZE</span><span class="p">(</span><span class="n">events</span><span class="p">);</span><span class="w"> </span><span class="n">i</span><span class="o">++</span><span class="p">)</span><span class="w"> </span><span class="p">{</span>
<span class="w">        </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">event</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyList_GET_ITEM</span><span class="p">(</span><span class="n">events</span><span class="p">,</span><span class="w"> </span><span class="n">i</span><span class="p">);</span>
<span class="w">        </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">previous</span><span class="w"> </span><span class="o">&amp;&amp;</span><span class="w"> </span><span class="n">PyObject_RichCompareBool</span><span class="p">(</span><span class="n">previous</span><span class="p">,</span><span class="w"> </span><span class="n">event</span><span class="p">,</span><span class="w"> </span><span class="n">Py_EQ</span><span class="p">)</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="mi">1</span><span class="p">)</span><span class="w"> </span><span class="k">continue</span><span class="p">;</span>
<span class="w">        </span><span class="n">previous</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">event</span><span class="p">;</span>
<span class="w">        </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">token</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">layout_token</span><span class="p">(</span><span class="n">it</span><span class="p">,</span><span class="w"> </span><span class="n">event</span><span class="p">);</span>
<span class="w">        </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">token</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="nb">NULL</span><span class="p">)</span><span class="w"> </span><span class="k">goto</span><span class="w"> </span><span class="n">error</span><span class="p">;</span>
<span class="w">        </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">token</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">Py_None</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">token</span><span class="p">);</span><span class="w"> </span><span class="k">continue</span><span class="p">;</span><span class="w"> </span><span class="p">}</span>
<span class="w">        </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">pos</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyTuple_GET_ITEM</span><span class="p">(</span><span class="n">token</span><span class="p">,</span><span class="w"> </span><span class="mi">2</span><span class="p">);</span>
<span class="w">        </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">PyObject_RichCompareBool</span><span class="p">(</span><span class="n">pos</span><span class="p">,</span><span class="w"> </span><span class="n">first_pos</span><span class="p">,</span><span class="w"> </span><span class="n">Py_LE</span><span class="p">)</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="mi">1</span><span class="w"> </span><span class="o">||</span>
<span class="w">            </span><span class="n">PyObject_RichCompareBool</span><span class="p">(</span><span class="n">pos</span><span class="p">,</span><span class="w"> </span><span class="n">last_pos</span><span class="p">,</span><span class="w"> </span><span class="n">Py_GT</span><span class="p">)</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="mi">1</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">token</span><span class="p">);</span><span class="w"> </span><span class="k">continue</span><span class="p">;</span><span class="w"> </span><span class="p">}</span>
<span class="w">        </span><span class="k">while</span><span class="w"> </span><span class="p">(</span><span class="n">index</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="n">PyList_GET_SIZE</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">))</span><span class="w"> </span><span class="p">{</span>
<span class="w">            </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">old</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyList_GET_ITEM</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">,</span><span class="w"> </span><span class="n">index</span><span class="p">);</span>
<span class="w">            </span><span class="kt">int</span><span class="w"> </span><span class="n">cmp</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyObject_RichCompareBool</span><span class="p">(</span><span class="n">PyTuple_GET_ITEM</span><span class="p">(</span><span class="n">old</span><span class="p">,</span><span class="w"> </span><span class="mi">2</span><span class="p">),</span><span class="w"> </span><span class="n">pos</span><span class="p">,</span><span class="w"> </span><span class="n">Py_LT</span><span class="p">);</span>
<span class="w">            </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">cmp</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">token</span><span class="p">);</span><span class="w"> </span><span class="k">goto</span><span class="w"> </span><span class="n">error</span><span class="p">;</span><span class="w"> </span><span class="p">}</span>
<span class="w">            </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="o">!</span><span class="n">cmp</span><span class="p">)</span><span class="w"> </span><span class="k">break</span><span class="p">;</span>
<span class="w">            </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">PyList_Append</span><span class="p">(</span><span class="n">result</span><span class="p">,</span><span class="w"> </span><span class="n">old</span><span class="p">)</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">token</span><span class="p">);</span><span class="w"> </span><span class="k">goto</span><span class="w"> </span><span class="n">error</span><span class="p">;</span><span class="w"> </span><span class="p">}</span>
<span class="w">            </span><span class="n">index</span><span class="o">++</span><span class="p">;</span>
<span class="w">        </span><span class="p">}</span>
<span class="w">        </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">index</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="n">PyList_GET_SIZE</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">))</span><span class="w"> </span><span class="p">{</span>
<span class="w">            </span><span class="n">PyObject</span><span class="w"> </span><span class="o">*</span><span class="n">old</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyList_GET_ITEM</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">,</span><span class="w"> </span><span class="n">index</span><span class="p">);</span>
<span class="w">            </span><span class="kt">long</span><span class="w"> </span><span class="n">kind</span><span class="w"> </span><span class="o">=</span><span class="w"> </span><span class="n">PyLong_AsLong</span><span class="p">(</span><span class="n">PyTuple_GET_ITEM</span><span class="p">(</span><span class="n">old</span><span class="p">,</span><span class="w"> </span><span class="mi">0</span><span class="p">));</span>
<span class="w">            </span><span class="k">if</span><span class="w"> </span><span class="p">((</span><span class="n">kind</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">NL</span><span class="w"> </span><span class="o">||</span><span class="w"> </span><span class="n">kind</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">NEWLINE</span><span class="w"> </span><span class="o">||</span><span class="w"> </span><span class="n">kind</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">INDENT</span><span class="w"> </span><span class="o">||</span><span class="w"> </span><span class="n">kind</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="n">DEDENT</span><span class="p">)</span><span class="w"> </span><span class="o">&amp;&amp;</span>
<span class="w">                </span><span class="n">PyObject_RichCompareBool</span><span class="p">(</span><span class="n">PyTuple_GET_ITEM</span><span class="p">(</span><span class="n">old</span><span class="p">,</span><span class="w"> </span><span class="mi">2</span><span class="p">),</span><span class="w"> </span><span class="n">pos</span><span class="p">,</span><span class="w"> </span><span class="n">Py_EQ</span><span class="p">)</span><span class="w"> </span><span class="o">==</span><span class="w"> </span><span class="mi">1</span><span class="p">)</span><span class="w"> </span><span class="n">index</span><span class="o">++</span><span class="p">;</span>
<span class="w">        </span><span class="p">}</span>
<span class="w">        </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">PyList_Append</span><span class="p">(</span><span class="n">result</span><span class="p">,</span><span class="w"> </span><span class="n">token</span><span class="p">)</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">token</span><span class="p">);</span><span class="w"> </span><span class="k">goto</span><span class="w"> </span><span class="n">error</span><span class="p">;</span><span class="w"> </span><span class="p">}</span>
<span class="w">        </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">token</span><span class="p">);</span>
<span class="w">    </span><span class="p">}</span>
<span class="w">    </span><span class="k">for</span><span class="w"> </span><span class="p">(;</span><span class="w"> </span><span class="n">index</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="n">PyList_GET_SIZE</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">);</span><span class="w"> </span><span class="n">index</span><span class="o">++</span><span class="p">)</span><span class="w"> </span><span class="p">{</span>
<span class="w">        </span><span class="k">if</span><span class="w"> </span><span class="p">(</span><span class="n">PyList_Append</span><span class="p">(</span><span class="n">result</span><span class="p">,</span><span class="w"> </span><span class="n">PyList_GET_ITEM</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">,</span><span class="w"> </span><span class="n">index</span><span class="p">))</span><span class="w"> </span><span class="o">&lt;</span><span class="w"> </span><span class="mi">0</span><span class="p">)</span><span class="w"> </span><span class="k">goto</span><span class="w"> </span><span class="n">error</span><span class="p">;</span>
<span class="w">    </span><span class="p">}</span>
<span class="w">    </span><span class="n">Py_SETREF</span><span class="p">(</span><span class="n">it</span><span class="o">-&gt;</span><span class="n">pending</span><span class="p">,</span><span class="w"> </span><span class="n">result</span><span class="p">);</span>
<span class="w">    </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">events</span><span class="p">);</span>
<span class="w">    </span><span class="k">return</span><span class="w"> </span><span class="mi">0</span><span class="p">;</span>
<span class="nl">error</span><span class="p">:</span>
<span class="w">    </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">events</span><span class="p">);</span>
<span class="w">    </span><span class="n">Py_DECREF</span><span class="p">(</span><span class="n">result</span><span class="p">);</span>
<span class="w">    </span><span class="k">return</span><span class="w"> </span><span class="mi">-1</span><span class="p">;</span>
<span class="p">}</span>
</pre></div>
</details>
<p>The failure case here seems somewhat obvious: the model is trained for token
efficiency for tool calling which also looks like code, and sometimes it seems to
be taking that code into a place where it should not be: the codebase.</p>
<h2>35 Hours on a Single Prompt</h2>
<p>I&#8217;m not really sure what to say here, but the slop machine was running for 35
hours until I turned it off.  In that time it produced a net addition of 75k
lines of code and it did not stop.  In the 35 hours it burned around 1B tokens
for a total of around 1200 USD in raw API costs.  It managed to produce 79 commits,
and that comes to a cost of around 15.5 USD per commit, and the agents exchanged
around 1400 messages.</p>
<p>I honestly do not need an agent to run for 35 hours on a single prompt.  It clearly
does not work or result in reasonable outputs.</p>
<p>So obviously: prompting it like this is stupid.  But when left unattended, it
<em>will</em> keep going, and earlier models did not do that.  Even Fable wasn&#8217;t as
crazy as that.  When you accidentally give it slightly too big of a task, it will
continue until it succeeds, even if it burns through an entire subscription.</p>
<p>And that&#8217;s more or less why right now I do not manage to trust this model much.
It has shown that it will commit slop, and it requires me to review it more as a
result.  Even if the failure rate is quite low, I would not want this.</p>
<h2>Disposable Code vs Committed Code</h2>
<p>In a world where code for tool calls is optimized for token efficiency and
&#8220;getting the job done&#8221;, I wonder if there is really enough signal going to the
training processes for &#8220;a human understands what is going on&#8221;.  I would say that
quite a lot of the code I get out of Astra is in my mind &#8220;objectively bad&#8221;.  But
it&#8217;s objectively bad by my human sense.  Maybe it&#8217;s objectively good for a
codebase that is entirely written by agents and only needs to be understood by
agents.</p>
<p>Which is why I&#8217;m honestly asking myself more and more why we are doing this.
These new models are absolutely amazing, for sure.  But I&#8217;m more and more
skeptical that the trajectory they are on still lends itself to present-day
software engineering processes.  The reason why I&#8217;m asking why we are doing
this is because I felt like we achieved a pretty good spot for
software engineering with those models, and that is the part of the AI economy
where it was possible to show a positive return.  But for how much more Fable
costs, for how much more Astra costs, I do not feel like the results are there.</p>
<p>In fact, with Astra and Fable I feel like not only are the costs astronomical,
but the models are also just not for me as a software engineer.  And presumably
that&#8217;s because these models increasingly are for other people.  For lawyers, 3D
artists, mathematicians, whoever uses computer use, etc.</p>
<p>And potentially as a byproduct of enabling all of this, you can now slop your
way to a one-shot 3D game over the weekend which looks impressive.  And probably
you can now run a software factory for as long as you don&#8217;t care about the code.</p>
<p>I&#8217;m sure I will get used to this, but man this stuff is weird.</p>
<small>
<p><strong>Postscriptum:</strong> speaking of weird: how is it that these models, in a sandbox,
with supposedly no way to communicate with other agents, manage to find the <a href="https://collusion.wiki/">same
public wikis</a> as a scratch pad for agent communication?
Did they collude during training runs to remember resources on the internet
which might come in handy in the future?</p>
</small>
<div class="footnotes">
<ol>
<li id="fn-1">
<p>I should clarify that I have done experiments like this before.  Typically
they do not run this long and the agent leaves behind a maybe imperfect but
still digestible piece of software.<a href="#fnref-1" class="footnote">&#8617;</a></p></li>
</ol>
</div>
]]></description>
    </item>
    <item>
      <title>Latent Powers</title>
      <link>https://lucumr.pocoo.org/2026/9/5/latent-powers/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/9/5/latent-powers/</guid>
      <pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>A few weeks ago I felt like it would be fun to see if I can make one of those
cheap Chinese CarPlay dongles run something other than the stock firmware.  The
idea was that rather than just forwarding CarPlay, why not do something more
interesting with them?  They all work quite similarly: they act as bridges
between your car and the phone.  From there they deal with video and audio
streams and pass some other data through.  Most of them also bring up a custom
UI for pairing and have a web interface that your phone can reach for updates.</p>
<p>Long story short: I had a conversation with Fable and Sol via Pi about what
could be done with such a dongle or whether I should use a Raspberry Pi instead
if I wanted to do my own thing there.  I figured it might be quite fun to run my
own code while still allowing regular CarPlay to pass through.</p>
<p>Through working with the LLM I learned about
<a href="https://github.com/catplay-labs/catplay">CatPlay</a>, which is a Rust
reimplementation of the CarPlay protocol that can run on Carlinkit devices.  In
particular, it can run on the Carlinkit Mini Ultra, which I figured would be
easy enough to buy.  I do have a few CarPlay adapters around, but I did not have
that particular model, so I bought one on Amazon.  Twenty-four hours later, I
had a device in my hand that was branded as a Carlinkit Mini Ultra, but instead
of being the Ingenic device that the original author used, it turned out to be
something else.</p>
<p>This is normally where the story would stop.  However, it&#8217;s 2026.  Armed with a
bit of knowledge about how these systems work, I managed to have some fruitful
discussions with Kimi K3 and Sol and figure out how <a href="https://github.com/catplay-labs/catplay-firmware/issues/4">flash the
device</a> and in turn,
how to make CatPlay compile for that SoC.</p>
<p>I guess that hacking these USB devices is not necessarily hard, but it&#8217;s
laborious and you can easily end up bricking your devices.  It also just sucks
because sometimes you need to work with someone else&#8217;s code that does not itself
run on your machine.  In the past, I would abandon many such projects for lack
of tenacity.  But my clanker is tenacious.</p>
<p>But so are <em>all of our clankers</em>.  Some of the projects we&#8217;re now attempting are
happening because of conversations we have with them.  In this case I did not
find or decide on CatPlay, the model did.  It was not the only suggestion, but
it became the best starting point after discarding others.</p>
<p>And I discover this more and more.  Particularly when we have solitary
interactions with these models, some of us &#8220;independently&#8221; decide to work on
similar projects.  When I talked with an acquaintance about CarPlay he also
mentioned recently that he decided to try something similar because he too
wanted to see if he can get his own agent be hooked up with the car.  And guess
what: he too learned about the CarPlay hacking community, and that it&#8217;s an
option, from the models and roughly around the same time.</p>
<p>It really got me thinking about how this could create situations in which
completely independent people end up building things they believe are their own
ideas.  Yet they were inspired or pushed towards doing something by a
conversation with an LLM — a conversation that someone else also had.  What if
we took paths, because those were the paths that were more likely with current
generation models?  There is a running joke in the AI builder community right
now that we&#8217;re all working on the same things, and in many ways it feels like we
are.  That might be because those things are obvious, or it might be partly
because we all use the same models with the same capabilities.</p>
<p>A few months ago, I first saw <a href="https://x.com/lucasmeijer">Lucas Meijer</a> share the
idea to make a model in Pi produce HTML reports rather than Markdown.  I thought
that was pretty unique.  Except, well turns out the models are probably trained
more and more for that (e.g. Claude Artifacts), and now it has become for many
the default choice for sharing reports.</p>
<p>How much of what we build comes from eliciting the same latent capabilities from
the same models?  Did the models make us prompt them that way?  Was it because
we shared ideas on Twitter and other communities that inspired us?  Or is it all
unrelated?</p>
<p>There is something powerful and strange about how LLMs diffuse knowledge and
capabilities, while perhaps also nudging us all simultaniously and independently
toward building the same things.</p>
]]></description>
    </item>
    <item>
      <title>Anger, Anxiety and Agency</title>
      <link>https://lucumr.pocoo.org/2026/8/24/anger-anxiety-agency/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/8/24/anger-anxiety-agency/</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>Sean Goedecke wrote a post arguing that <a href="https://www.seangoedecke.com/you-should-never-be-angry-at-work/">you should never be angry at
work</a> — a post
with which I strongly agree.  Anger can be a useful signal, but being angry at
work rarely improves the situation.  More often, it makes life worse for the
people around you, many of whom have no more power over the source of your anger
than you do.  I did learn that lesson, but it did not come naturally.  One thing
in particular that I learned is that in a company there is a shared vision, and
if you don&#8217;t agree with it and are not in a position to change it, you should
not start a mutiny, not even a small-scale one.  Nothing good comes from that.</p>
<p>In the discussion around that topic, one of the most upvoted comments on the
<a href="https://lobste.rs/s/mbmn1f/you_should_never_be_angry_at_work">Lobsters thread</a>
asked a question I had to think about quite a bit:</p>
<blockquote>
<p>How can you work in tech right now and <em>not</em> be angry?</p>
</blockquote>
<p>In the context of the thread, this was clearly also about AI and agents.  For
me, the emotions I would expect in tech vis-a-vis these new developments are
disorientation and anxiety, but not anger.</p>
<p>Anxiety as an emotion does not require someone to blame.  Right now, I find it
reasonable to feel anxious about an uncertain future.  Who knows what our
professions will turn into and what kind of world my kids will find themselves
in when they enter the workplace?  And if you&#8217;ve been in the industry for a long
time, will the skills you&#8217;ve spent years acquiring still matter?</p>
<p>But anger is different from anxiety because anger needs to be directed
somewhere.  The feeling of anger suggests that somebody or something is doing
something <em>to you</em>.</p>
<p>Who are you going to be angry at and why are you angry in the first place?  One
narrative that is pretty pervasive is that if AI will usher in productivity
gains, those gains are going to benefit companies rather than employees.  And
well at least <a href="https://thenextweb.com/news/meta-bosworth-ai-productivity-more-work-not-time-off">someone at Meta
wants
that</a>.
Yet I also find that plenty of people in leadership positions express doubt
about AI.  They see that an increasing share of their costs is being funneled
directly to some large AI labs.  They express worries about what will happen to
their data and whether these large companies will step into their space instead
of being partners.</p>
<p>My answer to the question of how you can not be angry in tech is that it&#8217;s by no
way the most only possible feeling.  First of all, instead of being angry, you
can simply be unsure.  The feeling of uncertainty is a much more productive
emotional state because it can lead to curiosity.  Even if you don&#8217;t find what&#8217;s
happening right now exciting, you can at least find it interesting.  We have
access to magic machines, and we can poke at them and see what happens.  The
second way is to feel genuine excitement.  Once you move beyond curiosity, you
can come away with a newfound feeling of power and freedom.  A lot of the gains
from AI aren&#8217;t turning into productivity gains that are reflected in company
profits but they&#8217;re showing up instead in the number of side projects shipped by
everybody not on their company&#8217;s time.</p>
<p>The fact that this is happening shows us that owners and founders don&#8217;t
necessarily know what will happen.  Ownership comes with agency, but it does not
provide foresight, and this change is disorienting for everybody.  I engage with
plenty of people who project confidence in public and are much less certain in
private.  Many of them are placing bets, but they are talking with confidence
about those bets, trying to keep their business afloat while the ground moves
under them.  They experience that uncertainty from a position where they can act
on it, and they are often standing somewhere with a megaphone to get others on
their side to improve their odds.</p>
<p>I feel that contradiction myself: I am simultaneously tremendously excited, but
I am also unsure what will happen next.  I do not know what it will mean to be a
programmer in the future, and, as the owner of a company, I am also not sure
where the high ground will be when this all settles.  Much of what I learned
over the years is changing rapidly, including ideas I considered fundamental to
my craft and business.  Some days that feels liberating, but on others I wake up
feeling like the ground is crumbling beneath me.</p>
<p>Anxiety is an uncomfortable emotion because it acknowledges that you do not know
what will happen and might not be able to stop it.  On the other hand, anger can
feel more actionable because, instead of saying &#8220;I don&#8217;t know,&#8221; you already have
someone to blame.  It turns a loss of control into a comforting story with a
villain.  But I feel that particularly when it comes to AI, it&#8217;s easy to pick
the wrong villain because of how disruptive the change is for everyone.  Your
engineering manager or leadership team might themselves feel uncertain about
their future and just try to bolster their own confidence by projecting clarity
and certainty.</p>
<p>That does not mean there are no villains.  When this all plays out, some will
profit and many will not.  I&#8217;m afraid we&#8217;re completely ignoring the impact this
has on society at large, the climate, and the balance of the world as a whole.
As excited as I am about the technology, I worry about Europe&#8217;s lack of ambition
and growing dependence on other countries.  I have a lot of complex thoughts
about what we&#8217;re doing as an industry right now.</p>
<p>I don&#8217;t know what the future of this industry will look like, and I don&#8217;t know
who will benefit from it and I don&#8217;t think I&#8217;m alone with that.  However I can
only urge anyone who feels anger and looks for a villain right to instead remain
curious instead.  To be curious enough to understand what is changing, excited
enough to experiment with it.  And then, from what we learn, earn the right to
decide when resistance is warranted and where to direct it.</p>
]]></description>
    </item>
    <item>
      <title>Fast and Hard Code</title>
      <link>https://lucumr.pocoo.org/2026/8/22/fast-hard-code/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/8/22/fast-hard-code/</guid>
      <pubDate>Sat, 22 Aug 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>One of the memes on Twitter is that &#8220;programming is solved now.&#8221;  I&#8217;m not sure
to what degree it is, but one thing is pretty clear: the act of familiarizing
yourself with a language no longer matters and some of the friction that
mattered for humans does not matter for agents.</p>
<p>As a result, LLMs make language choice much less consequential than it used to
be.  If you don&#8217;t like the choice, you can seemingly rewrite it in another
language and you can make it pick a language that you, as a programmer, are
entirely unfamiliar with.</p>
<p>Which in turn means that people can, and do, choose based on the marketing of
languages much more.  As a long-term Rust programmer I found it quite
fascinating to see people now ship Rust code who previously might not have
chosen it.  I attribute at least one part of this to two recent vibe shifts:
there is a lot more talk about wanting fast software, and about LLMs being
exceptional at optimizing code without regressing behavior.</p>
<p>Folks like Mitchell Hashimoto, Charlie Marsh, Jarred Sumner, Daniel Lemire and
quite a few others always carried a certain level of obsession with fast and
performant software and they also all happen to be receptive to agents writing
code.  Maybe as a result, or unrelated others are now joining in.  That&#8217;s
because with things like
<a href="https://github.com/davebcn87/pi-autoresearch">autoresearch</a> you don&#8217;t even
necessarily need to know all the tricks: you just need to put an agent on it —
though knowledge greatly helps!</p>
<p>If you look around, there are plenty of projects that want to be fast and small,
and they increasingly pick &#8220;hard languages&#8221;.  And it&#8217;s not just Rust that is
benefiting.  Even Zig — despite the fact that the creators and parts of the core
community are pretty negative on the whole AI thing — is too.  For instance
Cloudflare&#8217;s new
<a href="https://blog.cloudflare.com/artifacts-git-for-agents-beta/">Artifacts</a> service
uses a pure-Zig Git-protocol engine, compiled to a roughly 100 KB WebAssembly
module and Vercel released <a href="https://github.com/vercel-labs/fx">fx</a>, a Zig coding
agent advertised to be small and fast.  From what I can tell, all these projects
are largely LLM-assisted.</p>
<p>But it&#8217;s not just people picking less common languages but also that they are
increasingly working with &#8220;much harder&#8221; technologies.  All of a sudden I have
seen people do some really impressive stuff with DWARF files, eBPF, custom
network drivers, custom crypto and really old computing hardware.  Many of these
things were previously off-limits for lots of developers.  In some cases (eg:
crypto) you were even pushed away because those things were intentionally
gatekept by the people in the know.</p>
<p>So maybe the world will have more slop, but it might also have more developers
in it, that want things to be fast and small.</p>
]]></description>
    </item>
    <item>
      <title>What Is Reasoning</title>
      <link>https://lucumr.pocoo.org/2026/8/19/what-is-reasoning/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/8/19/what-is-reasoning/</guid>
      <pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>A few weeks ago <a href="https://arxiv.org/html/2608.09867v1">a paper was shared</a> that
showed how to extract reasoning traces from closed-weight models.  Together
with online discussions about tricking models into leaking them, it made me
investigate it more out of curiosity.  Twitter seems full of half-truths and
confusion about how this works, so perhaps this helps some to understand what is
happening.</p>
<h2>Hiding Traces</h2>
<p>Reasoning traces are usually hidden from us.  <a href="https://earendil.com/posts/session-portability/">We have lamented
this</a>, but mostly have to
accept it.  Open-weight models thankfully reveal them, and from their behavior
you can see that their traces can be long and confusing.  This is probably a
good reason to separate them from what is normally shown to users.</p>
<p>At minimum, UIs need to detect them.  The industry has done a good job at making
reasoning traces sound special and exotic, but they really are just text: the
model is trained to emit its thinking into a scratchpad as part of its response,
before its final answer.</p>
<p>GPT-OSS&#8217;s Harmony response format makes this easy to see:</p>
<div class="highlight"><pre><span></span>&lt;|channel|&gt;analysis&lt;|message|&gt;
I need to work this out ...
&lt;|end|&gt;&lt;|start|&gt;assistant&lt;|channel|&gt;final&lt;|message|&gt;
The answer is ...
&lt;|return|&gt;
</pre></div>
<p>The markers are special tokens, but the reasoning between them uses &#8220;the same
text&#8221; as the final answer (just that GPT chain-of-thought text sounds really
funny).  When the model samples the <code>analysis</code> channel token, a parser routes
the following text into a separate stream exposed through the Responses API.
For closed models, presumably a simple model redacts and summarizes it.</p>
<h2>Reasoning Effort</h2>
<p>How much budget goes to reasoning?  Earlier APIs exposed reasoning token
budgets, making it seem like a property of the sampling process.  In reality,
reasoning effort is baked into the system prompt.  GPT-OSS puts this into the
system prompt:</p>
<div class="highlight"><pre><span></span>Reasoning: low
</pre></div>
<p>That&#8217;s it.  Training produces the resulting behavior, such as emitting the
token sequence that switches to the <code>analysis</code> channel.  This also explains why
changing the effort invalidates the KV cache.  I think closed GPT models call
reasoning effort &#8220;juice,&#8221; since you can ask most models how much juice they
have.</p>
<p>In <a href="https://github.com/antirez/ds4">DwarfStar</a> for DeepSeek with max reasoning
this is added to the system prompt:</p>
<div class="highlight"><pre><span></span>Reasoning Effort: Absolute maximum with no shortcuts permitted.
You MUST be very thorough in your thinking and comprehensively decompose the
problem to resolve the root cause, rigorously stress-testing your logic against
all potential paths, edge cases, and adversarial scenarios.
</pre></div>
<h2>Don&#8217;t Think</h2>
<p>The destination of reasoning tokens is therefore a learned convention: the
model is trained to keep scratch work out of the <code>final</code> channel.  Trick it into
thinking it is in that channel and it may leak tokens.  We have even seen older
models, when thinking is disabled, reason into the bash tool and echo their
thoughts to <code>/dev/null</code>.</p>
<p>So in some sense the only &#8220;special&#8221; behavior for some models is not to think.
That at times is done by &#8220;mechanically&#8221; removing the model&#8217;s usual ways to think.
In <a href="https://github.com/antirez/ds4">DwarfStar</a>, disabled thinking uses the
prefill <code>&lt;/think&gt;</code>, while enabled thinking uses <code>&lt;think&gt;</code>, which are the tokens
that close and start thinking.  GPT-OSS doesn&#8217;t prefill but lets the model
decide either way on its own.</p>
<p>But presumably, some inference APIs prefill the opening token when reasoning is
enabled, so the model never samples it itself and might prevent the sampling of
the reasoning token when disabled since it can be trivially detected.  This may
explain why a <a href="https://gist.github.com/mitsuhiko/0904a3d89741e8e3bcca1ca93ea076de">custom <code>think</code>
tool</a> can
trick models into putting some reasoning where it should not go — but only when
native reasoning is disabled.</p>
<details><summary><small>Fun fact: this blog post triggered safey checks</small></summary>
<p>Hilariously enough I was unable to use GPT 5.6 terra for spell and grammar checking
on this blog post because of safety filters.  Had to switch to Kimi.</p>
<img src="/static/gpt-5.6-terra-spell-check.png" alt="GPT-5.6-terra refusing to spell-check this blog post" style="width: 100%">
</details>
]]></description>
    </item>
    <item>
      <title>Codeberg Divides</title>
      <link>https://lucumr.pocoo.org/2026/7/24/codeberg-divides/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/7/24/codeberg-divides/</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>Codeberg recently <a href="https://lobste.rs/s/ax914v/protecting_our_floss_commons_from_llms">changed its
terms</a> to
exclude projects that are largely written with generative AI.  Considering I
<a href="/2026/4/28/before-github/">want GitHub to get some competition</a> I have thoughts
about this.</p>
<p>GitHub&#8217;s governance has never been democratic and there is plenty about the
platform that I dislike.  Yet I do not need my infrastructure to be democratic
but I need it to be predictable and reasonably neutral towards the Open Source
software hosted on it.  A democratic provider without a clear constitution can
be worse at those things than a corporation is.</p>
<p>Codeberg is entirely within its rights to do run the platform like they want.
It is a German association with members and a democratic process, and that
process produced a result.  But a democratic vote says nothing about whether the
decision is a good one, particularly for the people already depending on the
platform.  A majority can still decide that certain projects and people no
longer belong.</p>
<p>The new <a href="https://codeberg.org/Codeberg/org/src/branch/main/TermsOfUse.md">terms prohibit
projects</a> that
mostly consist of code written by generative AI tools.  That&#8217;s fine, but these
days I could not assign authorship percentages to my recent projects.  For me
this rule is quite vague and I would bet that it makes it hard to enforce.  In
practice my assumptoin is that <a href="/2026/4/11/the-center-has-a-bias/">the center will
leave</a>.</p>
<p>If anything a harsher line would probably be preferable.  If Codeberg wants no
LLM involvement, it should say so.  On the other hand if the objection is just
spam and abusive resource consumption, it should write rules for those instead.
Now it defers the details of the policy to moderators and the communit which
already draws a much harder boundary than the text does, judging by the tone of
the discussion around it.</p>
<p>It is a shame that the Open Source and Free Software communities are splitting
this deeply over LLMs and agents.  These systems have problems, but these tools
are also becoming part of how software is made.  The Open Source world needs to
figure out how to engage with that future.  More importantly, LLMs if used well,
should be welcome to the Open Source community as they can be used to reclaim
control and take power away from large corporations.</p>
<p>I want GitHub to face true competition in the Open Source space.  I would
particularly like some of it to come from associations rather than another large
entity.  As a European project, Codeberg naturally matters to me even more.  I
just wish Codeberg were more forward-looking here and be willing to host the
Open Source software of tomorrow and not only software approved by the community
today.</p>
]]></description>
    </item>
    <item>
      <title>The Tower Keeps Rising</title>
      <link>https://lucumr.pocoo.org/2026/7/13/the-tower-keeps-rising/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/7/13/the-tower-keeps-rising/</guid>
      <pubDate>Mon, 13 Jul 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>I feel that some vibecoded software changes somewhat randomly and unexpectedly.
That made me think about Bruegel&#8217;s <a href="https://en.wikipedia.org/wiki/The_Tower_of_Babel_(Bruegel)">&#8220;The Tower of
Babel&#8221;</a> which shows
an already quite chaotic depiction of the Tower of Babel.  The story is one of
pride and ambition and ultimately why people no longer speak the same language.
But it is also a story about the unity that makes technological progress work.</p>
<p>The text begins with a technology upgrade:</p>
<blockquote>
<p>And they said one to another, Go to, let us make brick, and burn them
thoroughly.  And they had brick for stone, and slime had they for morter.</p>
</blockquote>
<p>They use it for a civilizational project:</p>
<blockquote>
<p>let us build us a city and a tower, whose top may reach unto heaven</p>
</blockquote>
<p>But when God assesses the situation the bricks are not what concern him:</p>
<blockquote>
<p>the people is one, and they have all one language, […] and now nothing will be
restrained from them.<sup class="footnote-ref" id="fnref-1"><a href="#fn-1">1</a></sup></p>
</blockquote>
<p>They get their power through coordination which they have because they share a
language.  They can use this to coordination to combine their powers and build
something no one of them could build alone.  God does not take away the bricks
or their knowledge of how to make them but their ability to understand one
another.</p>
<p>With AI-assisted programming we should get better tools which lets us build more
ambitious software.  That is certainly true at the level of the individual and
without doubt a developer with an agent can change a codebase dramatically
quicker.  But large software projects have never been limited only by how
quickly an individual can produce code but they are limited by how well people
can coordinate their understanding of the system they are changing.</p>
<p>The shared language of a software project is the common understanding shared
among its developers.  This language is rarely written down in one place but it
lives in documentation and code.  It can also just be something that comes up in
code review or watercooler conversations or when one engineer has to explain a
change to someone else.  It can be about the architecture of the code, the
tradeoffs made or which invariants need to be upheld.</p>
<p>In the days before agents some of that shared understanding was maintained by
friction.  If I wanted to change someone&#8217;s storage layer, I usually had to read
their code and ask them questions.  Changing that code might have required
coordination with another team whose service depended on it.  Some of this
friction was useful as it forced communication.  It also was a good touchpoint
for both of us discover if we still understood how the system worked.  I had
plenty of experiences in my career where through conversations with my fellow
engineer we collectively had to understand again why the system we had worked as
it did before we did some changes to it.</p>
<p>But with agents I can ask an agent to add OAuth, you can ask one to add caching,
and somebody else can ask one to rebuild the database from first principles and
make the UI pink.  Each change can be reasonable in isolation but since it&#8217;s
frictionless, none of us necessarily has to talk to the others or familiarize
ourselves with the code we are changing.  The more we use agents, the less we
feel the pain as agents feel none of it, and a useful signal is gone.</p>
<p>When I look at some vibecoded scaled-up projects the codebases mirror the story
of Babel mostly because nobody needs to communicate.  Obviously nothing stops us
to talk to one another, but nothing forces us either.  Every developer has a
tireless machine that can explain a corner of the tower and make whatever local
modification they want.</p>
<p>Unlike in the bible though, in AI-assisted engineering, construction can
continue after shared understanding has already collapsed.  The complete lack of
an immediate failure is what makes it curious and a bit disorienting.  The tower
does not fall, it just keeps rising.</p>
<div class="footnotes">
<ol>
<li id="fn-1">
<p><a href="https://www.biblegateway.com/passage/?search=Genesis%2011%3A3-6&amp;version=KJV">Genesis 11:3-6, KJV</a>.<a href="#fnref-1" class="footnote">&#8617;</a></p></li>
</ol>
</div>
]]></description>
    </item>
    <item>
      <title>Better Models: Worse Tools</title>
      <link>https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/</link>
      <guid isPermaLink="true">https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/</guid>
      <pubDate>Sat, 04 Jul 2026 00:00:00 +0000</pubDate>
      <description><![CDATA[<p>A very strange <a href="https://github.com/earendil-works/pi/issues/6278">Pi issue</a>
sent me down a rabbit hole over the last two days.  The short version is that
newer Claude models sometimes call Pi&#8217;s edit tool with extra, invented fields in
the nested <code>edits[]</code> array.  And not Haiku or some small model: Opus 4.8.  The
edit itself is usually correct but the arguments do not match the schema as
the model invents made-up keys and Pi thus rejects the tool call and asks to
try again.</p>
<p>That alone is not too surprising as models emit malformed tool calls sometimes.
Particularly small ones.  What surprised me is that this is getting worse with
newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the
older models.  In other words, the SOTA models of the family are worse at this
specific tool schema than their older siblings.</p>
<p>In case you are curious about Fable: I intentionally did not test it because I
was not sure if the classifiers they are running might downgrade me to Opus
silently.</p>
<h2>Tool Calls Are Text</h2>
<p>If you have not spent too much time looking at LLM tool calling internals, the
important thing to understand is that tool calls are not magic and use some
rather crude in-band signalling.  The model receives a transcript, a system
prompt and a list of available tools.  The server munches that into a large
prompt with special marker tokens.  Because the model was trained and
reinforced on examples of that format, at some point during generation it emits
something that the API or client interprets as &#8220;call this tool with these
arguments&#8221;.</p>
<p>For a file edit tool, the intended invocation payload might say something like
this:</p>
<div class="highlight"><pre><span></span><span class="p">{</span>
<span class="w">  </span><span class="nt">&quot;path&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;some/file.py&quot;</span><span class="p">,</span>
<span class="w">  </span><span class="nt">&quot;edits&quot;</span><span class="p">:</span><span class="w"> </span><span class="p">[</span>
<span class="w">    </span><span class="p">{</span>
<span class="w">      </span><span class="nt">&quot;oldText&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;text to replace&quot;</span><span class="p">,</span>
<span class="w">      </span><span class="nt">&quot;newText&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;replacement text&quot;</span>
<span class="w">    </span><span class="p">}</span>
<span class="w">  </span><span class="p">]</span>
<span class="p">}</span>
</pre></div>
<p>A harness then validates the arguments, performs the edit, and feeds the result
back into the model.  If validation fails, the model sees an error and usually
tries again.</p>
<p>How exactly that formatting happens is not known for the Anthropic models, but
some people have gotten out &#8220;ANTML&#8221; markers and they at times do leak also into
public communications.  To the best of my knowledge, the call above would come
out serialized like this from the model:</p>
<div class="highlight"><pre><span></span><span class="nt">&lt;antml:function_calls&gt;</span>
<span class="w">  </span><span class="nt">&lt;antml:invoke</span><span class="w"> </span><span class="na">name=</span><span class="s">&quot;edit&quot;</span><span class="nt">&gt;</span>
<span class="w">    </span><span class="nt">&lt;antml:parameter</span><span class="w"> </span><span class="na">name=</span><span class="s">&quot;path&quot;</span><span class="nt">&gt;</span>some/file.py<span class="nt">&lt;/antml:parameter&gt;</span>
<span class="w">    </span><span class="nt">&lt;antml:parameter</span><span class="w"> </span><span class="na">name=</span><span class="s">&quot;edits&quot;</span><span class="nt">&gt;</span>
[
<span class="w">  </span>{
<span class="w">    </span>&quot;oldText&quot;:<span class="w"> </span>&quot;text<span class="w"> </span>to<span class="w"> </span>replace&quot;,
<span class="w">    </span>&quot;newText&quot;:<span class="w"> </span>&quot;replacement<span class="w"> </span>text&quot;
<span class="w">  </span>}
]
<span class="w">    </span><span class="nt">&lt;/antml:parameter&gt;</span>
<span class="w">  </span><span class="nt">&lt;/antml:invoke&gt;</span>
<span class="nt">&lt;/antml:function_calls&gt;</span>
</pre></div>
<p>An important thing to note here is that this thing, while looking like XML, is
not really XML.  It&#8217;s just a thing they found convenient to tokenize and train
on.  The other thing to note is that a basic top-level string parameter appears
in-line whereas an array of objects is implemented via JSON serialization.
While I&#8217;m not <em>entirely sure</em> that this is how it works, there are some
indications that this is not too far off.  This will become relevant later.</p>
<p>There are two very different ways to make the model produce a structure like
this:</p>
<ol>
<li>You can <em>ask</em> the model to produce valid JSON matching a schema and then
validate it afterwards.</li>
<li>You can constrain the sampler so that invalid JSON, or even invalid schema
shapes, cannot be sampled in the first place.</li>
</ol>
<p>The second approach is what people usually refer to as grammar-aware or
constrained decoding.  The sampler masks out tokens that would violate the
grammar.  If the model is currently inside a JSON object and the schema says
only <code>oldText</code> and <code>newText</code> are allowed, the sampler can prevent it from
emitting <code>&quot;in_file&quot;</code> or <code>&quot;type&quot;</code>.  Grammar-aware decoding can be used both to
constrain something to be syntactically valid JSON and also to enforce specific
enum values or keys.</p>
<p>Without any form of constraints the model is merely following a learned
convention.</p>
<h2>The Failure</h2>
<p>Pi&#8217;s edit tool supports multiple exact string replacements in one call.  That is
why the arguments contain an <code>edits</code> array.  In the failing cases the model
produces entries like this:</p>
<div class="highlight"><pre><span></span><span class="p">{</span>
<span class="w">  </span><span class="nt">&quot;oldText&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;...&quot;</span><span class="p">,</span>
<span class="w">  </span><span class="nt">&quot;newText&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;...&quot;</span><span class="p">,</span>
<span class="w">  </span><span class="nt">&quot;requireUnique&quot;</span><span class="p">:</span><span class="w"> </span><span class="kc">true</span>
<span class="p">}</span>
</pre></div>
<p>or this:</p>
<div class="highlight"><pre><span></span><span class="p">{</span>
<span class="w">  </span><span class="nt">&quot;oldText&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;...&quot;</span><span class="p">,</span>
<span class="w">  </span><span class="nt">&quot;newText&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;...&quot;</span><span class="p">,</span>
<span class="w">  </span><span class="nt">&quot;oldText2&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;&quot;</span><span class="p">,</span>
<span class="w">  </span><span class="nt">&quot;newText2&quot;</span><span class="p">:</span><span class="w"> </span><span class="s2">&quot;&quot;</span>
<span class="p">}</span>
</pre></div>
<p>Across repeated trials I saw a whole zoo of invented trailing keys: <code>type</code>,
<code>id</code>, <code>kind</code>, <code>unique</code>, <code>requireUnique</code>, <code>matchCase</code>, <code>in_file</code>,
<code>forceMatchCount</code>, <code>children</code>, <code>notes</code>, <code>cost</code>, <code>oldText2</code>, <code>newText2</code>,
<code>oldText_2</code>, <code>newText_2</code>, and even an <code>event.0.additionalProperties</code> key inside
the edit object itself.</p>
<p>The most annoying part is that the actual <code>oldText</code> and <code>newText</code> payloads were
byte-correct in the invalid calls I inspected.  The model had in fact produced
the right invocation but then added nonsense at the end of the object.</p>
<p>The failure is also heavily context-dependent.  A fresh single-turn prompt like
&#8220;edit this file&#8221; did not reproduce it at all for me.  An agentic history where the
model had read files, diagnosed a problem and then composed a multi-line edit
could reproduce it.  And more annoyingly, not all transcripts will show that behavior.
In fact, I needed <a href="https://github.com/pasky">Petr Baudis</a>&#8216;s transcripts to
reproduce this for me at all!  In that user&#8217;s session continuing the session
caused Opus 4.8 to fail around 20% of the time.  Stripping thinking blocks from
history reduced the failure rate by half.  Turning on strict tool invocation
eliminated it in my runs.</p>
<h2>Why It&#8217;s Getting Worse</h2>
<p>My strongest hypothesis is that this is not random deterioration but a training
artifact.</p>
<p>When older Anthropic models were trained, they were trained on some tools (some of
which were documented).  But that training did not yet have a user-shipped
harness like Claude Code as the obvious target.  Modern Anthropic models are
most likely different because their post-training includes Claude Code or a
harness that looks very similar.  The model learns what a successful tool call
looks like in that environment.  It also learns what mistakes are tolerated by that environment.</p>
<p>Claude Code&#8217;s own tools are comparatively flat.  The ordinary edit tool is not
Pi&#8217;s nested <code>edits[]</code> shape; it is closer to <code>file_path</code>, <code>old_string</code>,
<code>new_string</code>, and an optional flag (<code>replace_all</code>).  Looking at Claude Code&#8217;s
client is very instructive: it contains retry paths for malformed tool use,
parameter aliases, type coercions, Unicode repairs and filtering of unknown
keys.  In other words, Anthropic&#8217;s own client appears to expect and accept a
fair amount of slop and repairs it, mostly silently.</p>
<p>If reinforcement learning happens in a harness like that, or a simulation of
one, then slightly malformed tool calls can still complete the task and receive
reward.  The harness fully absorbs the error and there is little gradient
against inventing an alias, adding a stray field or using a nearby parameter
name.</p>
<p>Worse, the model may become very strongly adapted to the canonical Claude Code
edit tool shape.  A different harness can present a tool with the same semantic
intent but a different schema.  Such a tool can increasingly be
off-distribution.  The better-trained model might actually fight you harder
because its prior is stronger.</p>
<p>This is not too surprising, but it is a change from how this was a few months ago.
When Opus 4.5 launched, it adapted to other edit tools exceptionally well.  In
fact, I was pretty convinced that we&#8217;re on a good path where the models are
more likely to adapt to any sort of tool shape that comes around for as long
as the instructions are good.</p>
<p>Now I&#8217;m somewhat worried about the track we&#8217;re on here.  Alternative tool
schemas might not just be unfamiliar.  They might be implicitly punished by
post-training that optimizes for one particular, forgiving tool ecology.  And
that ecology is not documented.  While there is a <a href="https://platform.claude.com/docs/en/agents-and-tools/tool-use/text-editor-tool">text editor
tool</a>
that is documented, you will see that this format is in fact not followed by
Claude Code.  What Claude Code does internally (which is a closed-source
harness) is hidden from you.</p>
<h2>The Slop Harness</h2>
<p>Claude Code is obviously closed-source but we can look at the minified code and
get some idea of what it does.  And honestly, it&#8217;s very forgiving of incoming
data.</p>
<p>For a start, Claude Code checks the model&#8217;s visible text for leaked <code>&lt;invoke</code>
markup.  It also emits some telemetry when that happens and then it has its
own state machine to retry such bad calls by pushing back to the model.</p>
<p>It has explicit Unicode escape repair which fixes broken <code>\uXXXX</code> sequences and
lone surrogates in string values.  It also has per-tool aliases for parameters.
For instance, <code>Edit</code> accepts <code>old_str</code> (presumably from the times when the models
were trained on the officially documented text editor tool), the newer <code>old_string</code>
from the schema, <code>new_str</code>/<code>new_string</code>, <code>path</code> as an alias for <code>file_path</code>, and some more.</p>
<p>It also silently filters out unexpected keys and it does not use <code>strict</code> mode
either.  The issue with <code>strict</code> mode is that Anthropic applies complexity
limits to the tool definitions that cause API requests to fail, so presumably
that&#8217;s why Claude Code does not attempt to use it.</p>
<h2>Strictness</h2>
<p>Will this problem be with us in other harnesses too?  One huge issue with
Anthropic is that the models are completely closed, and so is the harness.
Codex models are also closed, but at least the harness is not.  We also have
<a href="https://github.com/openai/gpt-oss">gpt-oss</a> which is at least a bit
interesting.  The models are explicitly trained to use OpenAI&#8217;s
<a href="https://github.com/openai/harmony">harmony</a> response format and there is
a lot of documentation that at least tells us how OpenAI people think about
this.</p>
<p>Harmony makes channels and tool-call content types part of the prompt format.  A
function call can look like this:</p>
<div class="highlight"><pre><span></span>&lt;|start|&gt;assistant&lt;|channel|&gt;commentary to=functions.get_weather
&lt;|constrain|&gt;json&lt;|message|&gt;{&quot;location&quot;:&quot;San Francisco&quot;}&lt;|call|&gt;
</pre></div>
<p>The important bit is <code>&lt;|constrain|&gt;json</code>.  The model can express in-band that
this message body is JSON, and an inference stack can use that boundary to
switch into JSON-constrained sampling for the body of the tool call.  Presumably
a bit of this also happens in Anthropic&#8217;s models, at least in <code>strict</code> mode
I would imagine.</p>
<p>The marker in harmony helps the sampler to detect when it needs to sample with a
specific grammar, and because it is part of the transcript, it makes that rather
easy to do.  For hosted GPT models, there is also an option to provide a
<a href="https://lark-parser.readthedocs.io/en/latest/grammar.html">LARK</a> grammar for
custom tools that need to adhere to something like this.</p>
<p>Anthropic appears different from that, though maybe not entirely.  If an array
of objects is represented as JSON, as it appears to be, then the model has to
write JSON inside the tool parameter.  There is probably basic
grammar-constrained sampling going on, and that may partly explain the extra
keys.  For a nested array parameter, that JSON includes escaped multi-line file
content inside string literals, inside one tag.  The unexpected,
made-up keys appear exactly at the highest-entropy point of that task: after
closing a several-hundred-token escaped <code>newText</code> string, where the model must
decide <code>}</code> vs <code>, &quot;...&quot;</code>.</p>
<p>Opus 4.8 and Sonnet 5 seem to have much stronger priors about what an edit tool
call should look like and that prior appears to be Claude Code&#8217;s edit schema: a
flat old/new string pair, plus the optional <code>replace_all</code> flag.  My guess is
that Opus has learned that an edit operation may have one extra optional field,
but under Pi&#8217;s nested <code>oldText</code>/<code>newText</code> shape it has no trained name for that
field.  So it samples a plausible name fresh each time, which is why the
failures produce dozens of random keys rather than one stable alias.</p>
<p>As <code>strict</code> mode in Anthropic appears to fix this, I presume that on the server
side they are refusing to sample a key that is not permitted by the JSON schema
structure.  That would also explain why they have limits to the complexity of
the tool definitions when strict mode is enabled.</p>
<p>So far, the Codex models I tested did not show this type of regression.  I tested
all available ones except 5.6, which I do not have access to yet.</p>
<h2>What This Means For Harnesses</h2>
<p>The uncomfortable lesson is that tool schemas are not neutral, at least not on
Anthropic models.  We like to pretend that a schema is an abstract contract and
the model is a general reasoner that will follow it, but that might no longer be
the case for some of the tools.</p>
<p>Tool schemas are somewhere in the distribution and some shapes are close to
what the model saw during post-training and some are far away.  Some are easy for
the provider&#8217;s hidden encoding (e.g. top-level attributes in ANTML), whereas some
require the model to write large escaped JSON objects inside nested arrays after
long multiline strings.  The model may be smart enough to understand the schema
and still be bad at sampling the exact shape under pressure.</p>
<p>If this type of model behavior continues, I wonder what the implications for
harnesses are.  Obviously one could turn on <code>strict</code> sampling in
Anthropic and the problem should go away.  On the other hand, that the model
has this behavior shows the impact that reinforcement learning has on them.
Fighting that prior is probably futile if you want to get the best model performance.</p>
<p>Right now the reality is that Claude Code is not open source and we cannot
really know what they are doing in their RL environments either.  We cannot assume
Claude-Code-trained behavior will transfer cleanly to your tools unless they are
a close match.  The more post-training happens inside one dominant harness, the
more every other harness will have to inherit its quirks.</p>
<p>I used to be more skeptical of strict grammar-constrained tool invocation
because constrained decoding can have quality tradeoffs.  I still think that can
be true in general, but this bug moved my priors significantly.  If the newest
models get better at solving the task while getting worse at faithfully emitting
an alternative tool schema, then the harness needs stronger guarantees
somewhere.</p>
<p>If you want to find out more, or you want to discuss this, consider reading the
<a href="https://github.com/earendil-works/pi/issues/6278">issue on the Pi tracker</a>.</p>
]]></description>
    </item>
  </channel>
</rss>