In January 2025, I had a fun hypothesis for a blog post : can LLMs write better code if you keep asking them to “write better code”? That was prior to the advent of robust agentic coding, but Claude Sonnet 3.5 was still able to iteratively improve on algorithmic Python code. The “better” instruction turned out to be underspecified: Sonnet abused that ambiguity to instead add a ton of useless features but the code was indeed faster. Even with the rise of agentic LLMs specifically RLHF ed to handle solving common pass/fail coding problems, optimization is generally not a part of that suite.

At the end of that blog post, I opined on a hypothetical future where LLMs could be able to write superfast Python code by instead writing Rust code and using PyO3 to bridge the languages to get Python’s ergonomics with Rust’s speed. An earlier draft of that post asserted that the same “write code better” instruction could instead be applied to the base Rust code and drastically improve its speed which would then propagate down to the Python code: however, back then I did not know enough about Rust and making such a claim would be too spicy without evidence.

After months of testing and experimenting since the release of Opus 4.5 made agentic coding more viable, I can confidently confirm that modern agentic LLMs can indeed write Rust code that is significantly faster than current state-of-the-art approaches if given appropriate guardrails and constraints . Additionally, as LLMs have made drastic improvements in coding in each successive frontier model release since Opus 4.5, the optimizations have become even better, cumulatively resulting in anywhere from 2x-20x speedup depending on the domain.

More importantly, this blog post is not a vaguepost and I am including both the prompts I used and the benchmark results. You’ve been warned.

At first, making software faster was a good quantitative way for me to test and compare these new agentic models. I used Rust as the target language primarily due to the Python integration and speed, but there are other aspects of the Rust language that make it particularly useful if a fast implementation is indeed discovered, such as memory safety and the ability to compile it to WebAssembly /WASM so it can run in a web browser without much effort. However, one important constraint I will follow that technically may not result in the fastest code is to forbid unsafe code whenever possible.

My first test case was reimplementing machine learning algorithms in Rust, which would lead to a meaningful productivity increase for me as a data scientist if I had faster and scalable tooling. At the time, it was arrogant to assume that I could beat battle-tested algorithms that have been iterated on for over a decade and are already written in C so Rust’s low-level benefits are not as pronounced. The algorithm I wanted to optimize the most was UMAP , which is a valuable algorithm for dimensionality reduction I used in my work, but scales poorly to big data and is very slow, with alternatives such as cuML being time-consuming to set up. UMAP Rust crates such as umap-rs already exist where I could just fork them and prompt Claude Opus 4.5 to add Python/PyO3 support, but as an experiment and learning experience I wanted to have the agent write the algorithm from scratch with minimal Rust dependencies in order to make optimizations at as low of a level as possible.

Rust has a comprehensive benchmarking tool with the criterion crate which all agents know how to leverage. criterion will run the benchmarks, track results across iterations to see if performance improved or regressed, and can calculate if this change is statistically significant or just noise.

Typical criterion output, depicting a 3.5x speedup relative to the previous run of the benchmark.

First, in the initial prompt for creating a Rust crate for UMAP, I asked Opus 4.5 to create benchmarks with different input data sizes since optimizations for small datasets may not work for large datasets and vice versa.

This approach created benchmarks using criterion and I manually reran the benchmark suite after prompting performance feature improvements such as using faer for faster linear algebra and using simsimd for faster SIMD operations . This quickly became cumbersome as I had to manually rerun each benchmark after each change to verify there are no speed regressions.

The way I prompt agentic LLMs is unusual: I typically provide the agents very long prompts prewritten in a Markdown document with the additional use of ALL CAPS and **bolding** for emphasis. This is to ensure I capture all nuances through the use of prompt engineering , along with several other tricks as detailed in this blog post. Although some may argue prompt engineering is dead as the latest models have become smart enough to correctly handle ambiguity, I strongly disagree as LLMs have also become much better at following said nuances.

Running a prompt from a Markdown document (right) by tagging the file (left), using Zed Agent .

After having enough confidence that the agent will not accidentally rm -rf the repo, I experimented with letting the agent be autonomous, giving them permission to iterate until they achieve a speed increase, hopefully.

It turned out “fast as it can be” is too ambiguous and Opus 4.5 was lazy so it tweaked a few hyperparameters without much of an actual speed increase and called it a day. What I needed was a clear target goal that can be pass/failed, so I refined the prompt:

This worked very well and not only did I get a 1.2x speed up on the benchmarks, but the agent continued after hitting the metric constraint and only stopped if a metric constraint was infeasible; in this instance, the agent hit 1.5x-2.0x speedups. The low-level Rust optimizations centered around a number of techniques including but not limited to: leveraging SIMD operations more aggressively, fusing functions, unrolling loops, creating intermediate caches, using Arc instead of borrowing wherever possible, and creating performance profiles based on input data (e.g. if the data is small, don’t use rayon data parallelism as the overhead erases gains).

I chose “1.2x faster” as a sanity test: if the goal is too high, the agent may cheat to achieve it through risky/verbose rewrites. Smaller changes are better since the agent can more easily isolate the cause of a speedup/regression, hence the note about iteration. After new frontier LLMs released such as GPT-5.3 Codex and Opus 4.6, I repeated this prompt unchanged for every new LLM and each were able to achieve a cumulative 1.5x-2.0x speedup over the previous pass. Going all the way to GPT-6 Astra over many months, that’s around 7.5x-32x faster than the initial implementation baseline.

This approach is hyperoptimizing for given benchmarks and therefore it could be considered benchmaxxing : a derogatory term for frontier LLMs that are only oriented to getting the high score on a benchmark which generalizes poorly to real-world use. However, if the benchmarks are sufficiently heterogeneous and truly representative of real-world use cases, then this is less of a concern. For this type of project, there are two ways to address concerns of benchmaxxing: 1) have the agent design diverse/unusual/adversarial input datasets instead of the generic “inputs up to 100000x768” and 2) enforce a quality gate on the output by comparing the output to a known correct implementation. In the case of machine learning algorithms, there is always a tradeoff between speed and quality, but in this instance it’s surprisingly easier to get the model fast, then make it correct. That is not how scientific engineering typically works, but it’s unlikely for a new implementation to match a known good implementation across many different benchmarks in an apples-to-apples comparison unless it’s truly correct.

An agent-optimized gradient boosted decision tree implementation which beats xgboost significantly in speed, but also sometimes quality! (MSE: lower is better; other metrics, higher is better)

Fortunately, there is a canonical implementation of UMAP with the Python package umap-learn and Python bindings to the Rust crate were already trivially added, so the new objective is simultaneous constraints: improve the code’s quality while capping the speed loss.

Indeed, the agentic Rust implementation had worse quality, but this followup prompt was successful and all quality metrics improved to near-parity with minimal speed loss. And this new crate was still 4x-15x faster than umap-learn with its Python bindings, and 2x-4x faster than the analogous Rust umap-rs implementation.

Results from the most up-to-date optimization pass for the Rust UMAP crate. In addition to faster speed, it matches or beats Python in most quality metrics.

Results from the most up-to-date optimization pass for the Rust UMAP crate. In addition to faster speed, it matches or beats Python in most quality metrics.

Convergence is found when an agentic iteration pass only results in a minor ~3-5% speed increase which may not be statistically significant while the agent adds a disproportionately large amount of code; the tradeoff is not worth it.

I ended up testing other machine learning algorithms with the same prompt progression: gradient-boosted decision trees (GBDT), multilayer perceptrons (MLP), graph networks, many of the typical algorithms from scikit-learn …and it worked on all of them . I don’t want to overfit on just optimizing machine learning despite that being ludicrously valuable in itself, so I employed a similar pipeline on more day-to-day software libraries to optimize them: templating engines, HTML parsing, and even web servers…and it worked on all of them once again.

These optimizations are not a simple process and you can’t just prompt the memetic “c’mon, try doing a breakthrough” to get better code because of the ambiguity of such a statement. I am not content with merely writing the fastest software: I want the software to be as fast as possible dammit. So, like my agent, I continued iterating and finding even more tricks to prompt engineer the agents into genuine breakthroughs.

All projects demoed within this blog post are in active development and results may not be indicative of their final releases…although I suspect they’ll be even better. 😇

It must be reiterated that agents can and will cheat if they can. In one example, I tested the agentic iteration pipeline on ballin —my 2D ball physics simulation in the terminal—in order to replace its rapier2d physics engine which was hitting a performance ceiling. Opus 4.5 was indeed able to speedup each physics step…a bit too well.

The numbers indicate the number of balls in the simulation: initially the sim lags at 15k.

The numbers indicate the number of balls in the simulation: initially the sim lags at 15k.

In headless_step , a 34,500x speedup and consistent performance across ball counts are both very very suspicious: upon manual inspection it turns out that Claude achieved the speedup by disabling the physics engine entirely . Which, fair play, but not ideal; a followup prompt did fix it and result in an overall performance boost (with added regression tests just in case).

For my experiments above, I used a custom Rust-oriented AGENTS.md; the most recent version of it is available here . Surprisingly, I haven’t had much of a need to update the core rules since my initial February agent experiments as LLMs keep improving at coding and I haven’t hit major issues that have necessitated additions. However, learning from my benchmark experiments, I added one more section to the AGENTS.md with some rules to mitigate sources of observed cheating:

With today’s agentic LLMs, these constraints have worked successfully, although I may still include them in the prompt as a force of habit. It’s easy to see if an algo gamed the benchmarks if you see the benchmark file in the git diff —agents can’t cheat that (easily, anyways).

Over time, I discovered a number of additional prompt engineering tricks and constraints that are also surprisingly successful in creating performance speedups.

In order to find the optimizations necessary to get 10x speedups over what’s currently state-of-the-art, the agents will need to think outside the box and avoid being anchored to what are currently best algorithmic practices. Therefore, I gave them both an explicit warning and commands of encouragement:

This worked, and resulted in a 1.2-1.5x cumulative speedup across benchmarks and different software domains.

Another trick I found to encourage agents to think outside the box is to invoke subagents. I hypothesized that forcing these subagents to research with different prompts could a) seed the parent agent with distinct ideas which could provide inspiration to the agent and b) serve as a check on the agent by reviewing distinct areas of the code for correct implementations. On that note I have a bone to pick with the software developer community: everyone talks about their army of subagent employees and how they’re amazing, but no one ever talks about how you invoke subagents within standard harnesses like Codex.

For difficult and highly parallel problems, the harness will automatically invoke a subagent tool to accomplish the work. However, one big problem is that in some harnesses, the subagent tool will invoke the subagent using the current size of the LLM, which can get expensive when using Opus/Sol-class models.

GPT 5.6 Sol subagents being invoked via the subagent Tool in Zed Agent. RIP my Codex quota.

I wanted to use a cheaper, small model like GPT-5.6 Luna for the subagents since they don’t need to write code, so I came up with a galaxy brain solution that works regardless of which parent harness is used: tell the model to run independent CLI commands that themselves invoke the agent:

This indeed works consistently; the constraints “long-duration”, “CLI command”, “do not use the subagent tool”, and “do not save their full transcript to a file” were added after the agent did inefficient things that wasted tokens. 7-12 is an arbitrary number range; since Luna is so cheap in usage, I chose a higher number than necessary. Not all subagents have salient ideas, but the parent harness can process and disregard bad ideas.

For my Rust word cloud crate, the parent GPT-6 Astra agent spins up Luna subagents with prompts addressing different areas of the codebase.

Overall, with subagent review, I managed to eke out another 1.2-1.5x cumulative speedup . Additionally, as of GPT-5.6 Sol, the “security” part of the prompt now works to provide ideas to harden the agent-generated code against unknown inputs while still getting the speed boosts.

With the style constraints enforced by my AGENTS.md, the code added with each optimization pass is reasonable at about 1k net lines of code (LoC) per commit. However, agents typically add the code to a single file and they will not proactively refactor. A bloated file is fine in development as long as it’s eventually fixed, so I wrote a prompt to perform said refactor:

I intentionally use SLoC (source lines of code) as the target metric instead of LoC because I don’t want the agent to remove comments to cheat said 20% removal.

Interestingly, this refactor is more computationally expensive than the actual coding and often takes longer. But it does eventually succeed, and during the benchmark pass to verify no severe regressions, the data there turned out to be unexpectedly weird:

Benchmark results after refactoring my graph network Rust crate.

Some benchmarks have a double-digit percentage speed increase / runtime reduction even though I didn’t explicitly ask the agent to optimize runtime speed. This doesn’t make intuitive sense for Rust as it’s a compiled language and with the constraints to follow all existing tests/functionality, it should compile to similar-performing code and not meaningfully faster code. 1 I’m certainly not complaining , though, so I added an additional constraint to at least avoid regressions:

Another useful approach is to create competitor benchmarks—that’s half the reason benchmarks are created in open-source software anyways. Let’s use templating engines as an example: Jinja2 in Python is one of the most famous packages in the language. In Rust, there are a few options: minijinja maintained by the same developer, tera inspired by Jinja2, and askama which differs from the previous in that it uses compile-time templates rather than runtime.

Therefore, after having Codex build a templating engine in Rust and run some optimization passes, I instructed Codex to build more benchmarks:

Immediately thereafter, I make a followup prompt with the same benchmarking techniques as usual, with one specific change:

Yes, I chose violence. And it worked , mostly.

S_J is the work-in-progress name for my template engine crate.

It got the 2x speedup against minijinja / tera in most benchmarks, more than typical agentic iteration alone. I don’t fully understand why: I was expecting it to inspect the code from other crates as a reference to research ideas for how to beat them, but it rarely does so. Perhaps agents have a competitive streak.

It did however lose against askama because of the compile-time difference. So naturally I told Codex to implement an additional compile-time path and then to beat askama .

Putting all the prompt engineering discoveries together thus far into a single prompt, I have created the Ur-Prompt for agentic iteration, available here . I encourage everyone to tweak the prompt for your use case and give it a try.

The last trick comes from a moment where I was frustrated that an Ur-Prompt pass resulted in zero improvement. So, with the failures primed in the session context and myself having the mindset of “things can’t get worse”, I tried a certain prompt.

I’m sorry. I’m so sorry.

…and it worked . It was able to achieve another 1.2-1.5x cumulative speedup over the already converged codebase. Since I made the prompt at the end of a session, the prompt is less inherently ambiguous: “don’t do what you already did thus far because it didn’t work well enough”.

In one case, I noticed that the agent just tweaked function hyperparameters to get the speedup, which is a valid breakthrough but not what I was going for. I took it to the logical conclusion by queueing an additional followup prompt:

This was enough to encourage the agents to fully try something different from either the Ur-Prompt pass or the first breakthrough pass, and it often achieved another 1.2-1.5x cumulative speedup on top of the previous speedup. It turns out that for software trained to follow user instructions, “just changing hyperparameters” is a grave insult that kicks the LLM into high gear.

When GPT-6 Astra released, to test it I did a sequence of Ur-Prompt + breakthrough + second breakthrough for all my repositories that had already converged with GPT-5.6 Sol, and Astra did indeed get the cumulative speedup, but in some cases it did find a real fundamental reimplementation of the algorithm that caused a 2x-3x speedup.

The result of running the breakthrough pipeline on my GBDT implementation; quality matched baseline. As of writing, I admit I don’t fully understand the breakthrough.