How to read a token-compression benchmark
Read token-compression claims through their denominators, paired tasks, cache costs and success criteria, with Tokenade versus RTK as a worked example.
Writing on AI-assisted development, engineering leadership, and open source.
Series: Context Engineering and The Real Cost of AI.
Read token-compression claims through their denominators, paired tasks, cache costs and success criteria, with Tokenade versus RTK as a worked example.
A verified $214 bill for a production rewrite, checked against the source transcript, sits beside a corpus of about 2,900 tech talks on coding-agent costs.
Claims that nobody can measure AI ROI and that 95% of pilots return zero are weaker than their usual retellings, once their sources are checked.
Stanford measures a 15-20% coding gain after rework. Stripe's reported 1,300 weekly pull requests show a different productivity signal.
Cloud FinOps put a cost study first and a named owner for the number second. Three practitioner cases show that sequence applied to token bills.
Opus 5.5 changes input, output and cache economics. Price the accepted task, including retries and review, before choosing a cheaper model.
A $589 billion single-day stock rout, an enterprise model platform, and a €25,000 backup server show three scales of vendor dependence.
Four conference talks and outside sources document Doctolib's AI adoption from 30 engineers to 600, without reducing it to one figure.
One gigawatt of AI data-center capacity costs about $50 billion. OpenAI has discussed adding up to 26, behind the series' cost figures.
Twenty days of my own Claude Code and Codex usage, measured three ways. The tools disagree by a factor of three, and dollars are not what binds.
Claude Code selected flow-lean but skipped its footer. A casing fix showed why installation, selection, and behavior need separate evidence.
Flow Lean fused three response skills. Its historical eval beat each comparator and caught a candidate offering a migration against an inferred target.
Near Melun in Seine-et-Marne, Campus IA plans 11 data centers. The Élysée calls it France's largest AI campus. Three buildings are authorized.
Près de Melun, Campus IA prévoit 11 centres de données en Seine-et-Marne. L’Élysée le présente comme le plus grand campus IA de France. Trois sont autorisés.
Authorizations, energy, impacts, and opposition: this dossier collects the figures, sources, and open questions around Campus IA in Fouju.
Autorisations, énergie, impacts, oppositions : ce dossier rassemble les chiffres, les sources et les questions ouvertes sur Campus IA à Fouju.
A release model for instructions, skills, Output Styles, hooks, MCP definitions and BM25 routing without confusing installed files with working behavior.
Two conference accounts describe CLAUDE.md bloat. How to scope project facts, procedures and response preferences, then verify loading and behavior.
Portable instructions require neutral sources, generated runtime outputs, release controls, and behavioral tests. Native primitives alone do not provide portability.
Personal CLAUDE.md to team AI instruction system for six engineers. How Méthode Aristote separates sources, shares modules, and catches behavioral drift in CI.
A map of context engineering responsibilities: architecture, specifications, agent identity and evaluation, with practical directions for existing skills.
Compare RTK, Tokenade, Headroom and prompt caching by mechanism, integration and evidence. Smaller tool output does not establish a cheaper completed task.
Use L0 to L5 to diagnose context drift, then maintain adherence through observation, repair, and replay instead of treating setup as finished.
Anthropic measured merged PRs; METR measured task duration in a different setting. What these results and long-context research establish, and what remains a hypothesis.
6 contributors in our git history. One is an AI. What the commit patterns actually look like after months of Claude Code, beyond the marketing claims.
The concepts I wish I’d known before week one: the agent loop, instruction scopes, context management, skill invocation, hooks, and client permission checks.
Choose between Claude Code and Cowork, then follow the Start, Build, or Scale path that matches your role and current operating problem.
Nine months of AI configuration in one production project: evolving responsibilities, profile-based generation and a measured reduction in always-on lines.
Everyone talks about shipping 10x faster with AI. Nobody talks about accumulating debt 10x faster. 7 months of production data from a real EdTech platform.
A non-developer modified 80 files in production in 10 days with Cursor AI. Exact timeline, AI config setup, and what it means for engineering teams in 2026.
Jeremy Twei coined it. Addy Osmani popularized it. Margaret-Anne Storey extended it to teams. Here's what I built to fight all three.
12 years scaling teams from 4 to 30+, then I quit the VP title. AI lets a solo builder match what small teams ship. Here's exactly what changed.
No posts match these filters.
Try another language or category.