TsumegoBench: How Good Are LLMs at Go?
I gave current LLMs simple life-and-death problems. More interesting than the ranking was how quickly I could build a benchmark of my own.
Read moreAI, personal tools, cost reality, small experiments, and the life around them.
I gave current LLMs simple life-and-death problems. More interesting than the ranking was how quickly I could build a benchmark of my own.
Read moreHow I leave my Mac at home and still access it from my iPad and work with coding agents.
Read moreAfter the Anthropic cut, I compare my Claude Code usage against API prices and look at which models still fit into my 100-dollar budget.
Read moreWhy I'm cancelling my Anthropic subscription, what made Claude Pro so important for my agents, and why we should get ready for real API pricing.
Read moreWhy I decided to collect one new experience for every letter in 2026, and how that turned into my ABC Challenge Tracker.
Read moreWhy I started another website
Read more