OpenCode Go’s quota is the price
I am a big fan of OpenCode Go. For $10 a month after the first month, it gives developers access to a useful set of coding models, and the price feels almost...
I am a big fan of OpenCode Go. For $10 a month after the first month, it gives developers access to a useful set of coding models, and the price feels almost...
I suspect the definition of personally identifiable information will become stricter as AI lowers the cost of piecing together an identity. A recent paper, L...
I keep coming back to Brad Setser’s work on Chinese economic data because he shows how much interpretation sits behind an apparently simple metric. In his la...
The report that GPT-6 Astra helped decode a 1918 German ADFGVX message is another reason I expect interesting historical research from AI. The result still n...
The early results on GPT-6 Luna make it look like a very strange model. OpenAI is presenting Luna as a cheaper, more efficient successor to GPT-5.6 Luna, whi...
I tested TypeSafe’s Jev on chance, hiring, and political judgement. The results are a useful warning about what low latency does not tell us.
Back in 2020, neural-network research often looked almost embarrassingly simple from the outside: increase the parameter count, add data and compute, and exp...
I like the idea of navigating by Earth’s gravity when GPS is unavailable. It has the obvious benefit of being passive and difficult to jam, and it is also a ...
I have been reading Morgan Linton’s VulcanBench, an open-source harness that tests coding agents on real software tasks rather than asking a model a sequence...
On 3 September, Eric Lu at Cognition factored RSA-260, an 862-bit challenge number, with a GPU-accelerated version of CADO-NFS and help from Devin. On 19 Sep...
A story circulating on X alleges that Z.ai’s ZCode desktop coding app quietly packages a user’s workspace, including much of its .git history, LFS objects an...
Google Research, working with HHMI Janelia and other collaborators, has released a complete wiring diagram of the male fruit fly’s brain and central nervous ...
I am quite interested in Devin Fusion, though less because of the implementation than because of the idea behind it. Fusion models and automatic routing seem...
Today I revised the explanation of my explosive condensation project. The paper, “Explosive condensation in symmetric mass transport models”, was largely my ...
I built a small game called Hong Kong Tycoon, set in the early colonial period of Hong Kong. The idea began with two games I played when I was younger. One w...
Last week, I wrote about what Gabriel’s horn can teach us about mathematics education. That post began with a strange mathematical object whose volume is fin...
At a mid-size company, AIOps starts with a reasonable stack and can quickly turn into a cupboard full of model credentials.
I wrote earlier that software companies should expose themselves to agents or risk being bypassed, and that API usage was the natural replacement for per-sea...
Simon Willison recently described a small experiment that says more about the direction of software than its subject. He installed Blender on a Mac, asked Co...
Nike’s recent struggle with direct-to-consumer retail is a useful reminder that product complexity is not always a problem to be removed.
I recently shared the paradox of Gabriel’s horn with my daughter. It is a nice mathematical object because the underlying idea is not especially difficult. T...
Two things surprised me about GPT-6 Astra. The first was the headline result on ARC-AGI-3, where OpenAI says Astra reached 99.9 percent. The second was the e...
The paper on S2TDM, a spatial-spectral transformer-based diffusion model for hyperspectral image denoising, fascinated me. I am a layman to the topic, so thi...
I am quite impressed by how recent and practical Stanford’s CS 329Z, Engineering AI Agents, is. The course covers compound AI systems, retrieval, tool use, a...
Metrics are being solved faster than I expected.
I have updated my Artificial Analysis graph to v4.2. What interests me is how quickly the index changes. Metrics used to be something models worked around, b...
We need metrics for the agent harness as well as for the models it wraps. Metrics already exist for parameterisable end-to-end processes, so if the harness i...
The complete male fruit-fly connectome is fascinating to look at, although I am unsure whether the visual beauty exceeds its practical use. A few neurons can...
Program as Weights is a brilliant idea. Fragmented, personalised software is expensive to fine-tune in the usual way, so compiling English specifications int...
Paul Graham’s advice is simple: “Make good new things.” He argues that making something is often the best evidence that we have thought clearly, whether the ...
I built an Elo explorer for every rikishi in the data set, from 1958 to the present. The first thing that bothered me was Taihō.
Thomson Reuters’ report on Thomson makes a strong claim: an institution can take an open-weight model, improve it through continual learning, and end up with...
GLM has made the headlines, but people often forget how hard the engineering still is.
The mud floods along the Nepal-China border are difficult to look at. Roads, homes, bridges and lives can disappear under a moving wall of water, mud and roc...
I recently spoke with an intern on the team about education. We began by comparing the subjects we had studied, and soon found ourselves talking about litera...
Recent research shows that agents can collaborate toward a common objective through a shared environment, without communicating directly (Anthropic). Elsewhe...
Brad Setser puts China’s true current account surplus at about $1.2 trillion, or 5.5% of GDP (his chart here). Officially it is 4%, over $600 billion across ...
Tencent compressed Hy4-preview, its open-weights 770B model, from 1.5TB to about 200GiB of GGUF, over seven times smaller. The scheme, MIX-STQ1_0, lets calib...
Someone ran the experiment I keep hoping to see. @superalesha spent 1,351 hours of rented Blackwell time on the same model, Qwen3.8-27B, in nine versions: th...
I gave my wife her own Hermes agent and set the default model to DeepSeek V4 Flash, one of the cheapest models you can run. She came back with a series of th...
Z.ai served GLM-5.3-Flash on Chinese chips this week and claims per-token cost on par with Nvidia, after a 3x serving improvement. But perhaps the most impor...
One night YC quietly gave its AI agent full access to the production database. The agent became 10x more useful. That experiment, told in the Lightcone episo...
Yi Fuxian blames the one-child policy for China’s fed-up youth in this Project Syndicate piece. The mechanism he describes, fewer children weakening househol...
I built an Elo rating for every rikishi who competed since 1958. The dashboard is live at yuxichau.com/projects/sumo-elo/.
Ethan Mollick got early access to Claude 5 Fable, the first Mythos-class model, and wrote up what it felt like to use. The facts alone are worth the read. He...
On August 20, a model called Ox Alpha appeared on OpenRouter with no company name, no press release, and no logo. Just a stealth label and a price tag of zer...
Orca Router dropped uncensored weights of Qwen3.8-27B on Hugging Face this week: full precision (gated), plus open GGUF, FP8, and MLX builds. The GGUFs pulle...
I read a post by Paul Graham on X a while back. It was mostly about something else, but somewhere in the middle he made a one-sentence statement about the fu...
Everyone who works in a financial institution has an opinion about governance, and most of them are cynical ones. If you come from a security background, gov...
Tyler Cowen wrote Seven Ways to Avoid Losing Your Job to AI for The Free Press. The short version: look for messy jobs that are hard to describe, be wary of ...
In January 2024, a team led by French archaeologist Stéphen Rostain published a landmark paper in Science: “Two thousand years of garden urbanism in the Uppe...
In most companies running a generative AI programme, I’m the person people send when they can’t decide which model to pick. The question comes around every f...
I love it when you find a bug in the world. Not a software bug, but an error in the official story of things. A detail that doesn’t quite fit.
LLMs are impressive. The real work begins after the demo, when you have to make an AI do something useful. That means giving it tools, connecting it to data,...
I’ve noticed something when using GitHub Copilot. If I start writing code to analyze data, it almost always suggests Python and the pandas library. Usually a...