Caveman started as a joke: make Claude talk like a caveman so my AI bill stopped

By Julius Brussee · October 2, 2026 · Curated by George's Blog

Caveman started as a joke: make Claude talk like a caveman so my AI bill stopped beating my grocery bill.

108,000 GitHub stars later, today it grows up.

Caveman 3.0 is fully Apache-2.0. The proxy, the engine, the CLI, the SDKs. All of it. Tribe hate BSL.

And it doesn't just speak like a caveman anymore. It reads like one.

Here's what I kept seeing: the expensive part of a coding agent isn't what it writes. It's what it reads. Every turn it re-reads 70 KB of logs, test output and JSON where one line matters and the rest is noise. Then you pay for that noise again next turn.

So Caveman now sits between your agent and the model and shrinks what it reads, before it's sent:

→ One 70 KB log file: 22,810 tokens → 348. The one FATAL line the agent needed is still there.

→ CSV, YAML, test output: 98–99% smaller, every answer field kept.

→ Whole Claude Code sessions: 33% fewer tokens on our benchmark, 18/18 answers still right.

→ The originals stay on your machine, and the agent can pull any of them back.

Plus caveman learn, which reads your own agent history locally and shows you where your tokens actually go. Mine: a 423-line CLAUDE.md riding along with every single message.

The lesson from building this: bigger context windows didn't fix anything. Agents don't need more room. They need less junk in it.

Building teams on agents? The middleware is now 1.0 for TypeScript and Python. Same compression, wrapped around the LLM call you already make.

Still building this alone, at 20. Next up: Caveman Cloud.

View the original post on LinkedIn