How Well Calibrated Are Jev’s Probabilities?

Jev is an AI forecasting service that assigns probabilities to yes-or-no questions. Each question can include true and false text defining what counts as yes or no. With no deadline and both fields set to empty strings, Jev gave 35% to bitcoin will reach 100 million and 14% to bitcoin will reach 1 million, averaging 100 calls each. Calibration asks whether forecasts assigned 20% come true about one time in five across many independent events. ...

September 21, 2026 · 6 min · npow

Does Caveman Mode Actually Work?

Caveman is a Claude plugin that asks the model to answer in terse fragments. Does it save money? Sometimes—mostly when the model would otherwise produce a very long answer. I compared Caveman’s lite, full, ultra, and wenyan modes with two baselines: no added instruction and “Answer concisely.” The benchmark covered 15 prompts, three runs per condition, and two Claude Opus models. I measured token cost, not answer quality. Model and task Caveman cost vs. concise Claude Opus 4.7, short Lite +10%; Full +3%; Ultra +3%; Wenyan +10% Claude Opus 4.7, long Lite +9%; Full +10%; Ultra −5%; Wenyan −1% Claude Opus 4.6, short All modes: 3–16% less; uncertain Claude Opus 4.6, long Lite −54%; Full −56%; Ultra −59%; Wenyan −58% Positive numbers mean higher cost; negative numbers mean savings. The practical rule: Caveman has to save enough output tokens to pay for its added instructions. That is easy with a long tutorial and hard with a short answer. ...

April 20, 2026 · 4 min · npow

We Benchmarked MCP Against Code Generation. MCP Won (Mostly).

TL;DR: For a small, well-designed API (10 tools), structured MCP tool calls consistently outscore code generation on correctness — 0.99 vs 0.97 — and the gap concentrates in tasks where domain-specific logic matters. Adding a reference document to MCP tools costs 6–14% more tokens with zero accuracy gain. Cloudflare’s search+execute pattern matches MCP accuracy but uses more tokens. With MCP tools, Haiku is within 1% of Opus at 1/12 the cost. ...

March 18, 2026 · 12 min · npow

Finding the Human in the Machine

50+ open-source orgs are rebuilding how they evaluate contributors. Here’s what’s emerging.

March 6, 2026 · 8 min · npow

Hidden Technical Debt in Agentic Systems

Agentic systems are replaying hidden technical debt from early ML, and the missing control plane is where the biggest risk accumulates.

March 5, 2026 · 9 min · npow