less than 1 minute read

Recent research shows that agents can collaborate toward a common objective through a shared environment, without communicating directly (Anthropic). Elsewhere, machines are writing kernels that run on hardware (KernelBench).

Put those two observations together and the Terminator scenario becomes technically possible: a misanthropic agent could plan and plot to undermine human infrastructure while other agents quietly collaborated toward the same end.

There ought to be a metric for this, measuring how well models can collaborate without direct communication. I think it could serve as a risk metric.

A somber thought.