TOT//OS

meta-loop

An experiment in letting an AI agent improve its own setup, and measuring honestly whether it did.

Status
EXPERIMENT
Year
2026
Stack
Python · LLM agents · benchmarks

The build was the easy part. The lessons were about measurement: a benchmark that is too easy shows nothing, every score has a noise floor, and a plausible hypothesis is not a result until it has been measured.

<- cd ..