Models, harnesses, frameworks and tools for people who ship AI agents
Leaderboards for the current stack, plus articles on how the pieces fit together. Every page carries a half-life meter showing its age and whether it is still worth acting on, because writing about AI tooling goes stale faster than anyone admits. Start with models, harnesses and frameworks.
Every meter below is computed live from its publish date and half-life. Nothing here is set by hand.
The half-life meter
Old posts mark themselves out of dateLatest writing
Continuous log-
AI agents are now a security threat: the PaperCut and RubyGems incidents
In May and August/September 2026, AI agents were used to breach hundreds of organisations in two separate campaigns. What happened, and what it means for agent safety.
-
Meta Muse, the Agents API and AIforce: the platform war is now a harness war
In four weeks Meta, OpenAI and Salesforce each started selling the layer around the model: a personal agent on its own VM, the Codex harness as an API, and an enterprise control plane. What shipped, when, and how to stay portable.
-
JEV: what TypeSafe's decision model actually does
A decision model, not a chat model: one typed, calibrated call in 70–500 milliseconds, output tokens free. What it is for, where it breaks, and why agents have already run 8.1 billion tokens through it.
-
Agentic AI workflows: what they are and how to build one this week
Agentic AI workflows let software make decisions and take action without waiting for you. How they differ from regular automation, and a step-by-step n8n template for building your first one this week.
-
When an agent rewrites its own instructions: /refine, skills and protected files
Prime Agent's /refine and Hermes's learned skills both let an agent change how it works. That is the point, and the risk. How each project limits the damage, and how to review what an agent teaches itself.
-
GPT-6 Astra needs less prompting than you think
OpenAI's guidance for GPT-6 Astra is clear: shorter skills, fewer approval gates, less hand-holding. The models got better, so your prompts can get smaller.
-
Harness engineering lessons from production AI chatbots
Lessons from deploying AI chatbots on client sites: prompt architecture, context management, fallback chains, cost control, and which harness patterns hold up in production.
-
Hermes Agent 0.21.2: the state.db release and the credential vault
After a big 0.21.0, Hermes spent its next two point releases on reliability. What broke, what the 11 September patch fixes, the new password-blind vault, and how to update safely.
-
Skills frameworks are everywhere. Here is what makes one worth using.
OpenAI, Hermes, HumanLayer and a dozen GitHub repos all shipped skills systems in the same month. The shared pattern is clear enough; the implementations are not. What separates a good skill from a bad one.
-
MCP security is an unsolved problem. Here are the four attack surfaces.
Three papers published in August and September 2026 found at least four distinct attack surfaces in every MCP deployment. What they found, and what to do about each one.
-
Sandboxes and self-writing agents: what isolation Prime Agent and Agent Zero give you
Both agents write and run their own code. One tells you outright that it is not a sandbox; the other wraps everything in a container. What each boundary protects, and what neither does.
-
Hermes, Agent Zero or OpenClaw: which self-hosted agent should you run?
Three open-source agents you can run yourself, built for three different jobs. A side-by-side comparison of where each lives, how it is isolated, and the kind of team it suits.
-
Engrim: local-first memory that survives context clears
Every time you clear your agent's context, months of architectural decisions vanish. Engrim keeps them in a local SQLite store that works across Claude Code, Cursor, Antigravity and Codex.
-
NLIP, A2A, MCP: the three protocols your agents actually need
Three protocols claim to solve agent interoperability, and each one works at a different layer. What they do, where they overlap, and what to build on today.
-
How we built a 24/7 AI enquiry assistant for static sites
Static sites cannot run server-side logic, but clients wanted an AI chatbot that could handle enquiries, qualify leads and send emails. The architecture, the trade-offs, and what worked in production.
-
Running Claude Code, Codex and Hermes through Ori: a working setup
Install Ori, sign in once, and launch your usual agents on any OpenRouter model. The exact commands, the model flag, the controls that carry over from your organisation, and where it still trips people up.
-
DaeOne: one interface to run your entire website
Website ownership is fragmented: analytics, email, SEO and support all live in different tools. DaeOne pulls everything into one AI-powered dashboard.
-
Local LLM inference in September 2026: what actually works
Running models on your own hardware is no longer a compromise. Where local inference stands right now: the tools, the hardware, and where the trade-offs land.
-
Agent Zero: giving an agent a whole Linux computer
Agent Zero's pitch is not a better tool list but a whole desktop in a container: a browser, an office suite, a terminal and isolated projects. What that makes possible, and the risks that come with it.
-
Hermes Agent 0.21: bot society, cron memory and steerable subagents
The 31 August release turns Hermes from a single assistant into a small society of named agents, gives scheduled jobs a memory, and lets you steer subagents mid-flight. Here is what changes for a small-business deployment.
-
Ori Harness: OpenRouter's wrapper that configures your agent for you
Ori launched on 4 August to run Claude Code, Codex, OpenCode and Hermes through OpenRouter without hand-editing a dozen environment variables. What it does, what it does not, and who it is for.
-
Prime Agent: a harness with one tool and a habit of rewriting itself
Prime Intellect's open-source harness hands the model a single persistent Python kernel and lets it edit its own skills. What that buys you, what the benchmark numbers do and don't say, and why it is not a sandbox.
-
Harness Engineering is now a discipline. Here is what that means.
The model is becoming commoditized. The harness — everything around the model — is where the engineering value lives. Three data points prove it, and the implications are immediate.
-
OpenRouter: the LLM gateway that lets you test 400+ models with one API key
OpenRouter gives you 409 models from 59 providers behind one API key. 18 free models including frontier-class options. The practical starting point for autonomous agents and AI frameworks.
-
NVIDIA Nemotron 3 Ultra: frontier reasoning, built-in safety, and completely free
NVIDIA's Nemotron 3 Ultra matches frontier models at 3-10x fewer active parameters, with safety guardrails baked into training, not bolted on. Free on OpenRouter.
-
AI Model Price-to-Performance Guide
A living comparison of 400+ models across five price tiers. Updated every Monday with fresh OpenRouter pricing, new releases, and benchmark data.
-
What Claude Fable 5 actually does differently
Anthropic's most capable widely released model runs with thinking always on and a 1M context. Here is what that changes about how you build on it.
-
Xiaomi MiMo: what you get for the money
A 1M-context omnimodal model at roughly a fiftieth of frontier pricing. Where that trade is worth making, and where it very much is not.
-
Anatomy of an agent harness
A model is not an agent. The gap between them is six pieces of unglamorous engineering, and almost every agent that fails in production fails in one of them.
-
OpenCode Go: $10 a month for eighteen models
A subscription that fronts open-weight coding models at bulk rates. The maths works out — provided you understand what the three spending caps actually do.
-
Nous Portal: one account instead of nine API keys
Nous Research's subscription bundles a model catalogue, a managed tool gateway and cloud hosting behind a single sign-in. What it replaces, and what it locks in.
-
Hermes Agent for small businesses
A self-improving agent that runs on a £5 VPS and answers on WhatsApp is a genuinely different proposition for a six-person company than a coding copilot is.
-
What to automate first with Hermes
Five jobs a small business can hand to an agent this month, ordered by how little damage they do when the agent gets it wrong.
-
Setting up Hermes Agent: a working guide
Install to first useful automation, including the parts the quickstart leaves out — where it runs, which model to point it at, and how to stop it doing something expensive.
-
An introduction to OpenClaw
The most-starred project on GitHub is a local-first personal agent with no subscription. What it is, why it spread, and what running it actually commits you to.
-
An introduction to Paperclip
An org chart for your agents: roles, reporting lines, budgets and audit trails. The orchestration layer above whichever coding agents you already run.
-
An introduction to CrewAI
Role-playing agents that divide work between them, plus an event-driven layer for when you need the sequence to be predictable. Knowing which half to use is the skill.
-
Context engineering replaced prompt engineering
Long context windows didn't end the problem, they moved it. What goes into the window, in what order, and what gets thrown away is now the job.
-
MCP in practice: what the protocol is actually for
The pitch is "USB-C for tools". The reality is more specific, and knowing where the boundary sits saves you from building a server you didn't need.
-
Pick a framework, or pick a loop
Most agent frameworks are a while-loop with opinions. Here is how to tell whether you need the opinions, and what you give up when you take them.
-
Evals that actually catch regressions
A scoreboard that only goes up is measuring the wrong thing. Building the small, mean test set that tells you when a change made things worse.
Why the meters exist
Documentation should say when it has gone out of date.
We check harnesses against the current provider APIs. When an SDK changes or breaks in a way that silently changes behaviour, the affected post's decay meter moves.