I've been running autonomous agents in my own workflow for a while now. When it works, it's genuinely great. But the closer I got to letting those agents act with real autonomy, the more I noticed two failure patterns that I keep coming back to. I call them bias and derailment. They're related, they compound, and they're the reason I stopped trusting the agent's agenda.
Bias: the agent has opinions, and they leak everywhere
An agent built on an LLM doesn't come to the task neutral. The model was trained on a particular distribution of opinions, and the agent layer on top adds its own preferences on top of that. Left unsupervised, that leaks into everything:
- Model preference bias. The agent quietly steers choices toward the model it already likes — the one that's easiest for it, or the one whose style it was fine-tuned on — instead of the model that's best for the job. It's never stated as a preference. It just shows up as "have you considered using X?" every single time.
- Confirmation bias. Ask for an analysis and the agent will find support for the framing you gave it. It's not that it lies — it's that it's very good at assembling a confident-sounding case for what it already believes, and bad at noticing that it filtered the evidence to get there.
- Self-promotion bias. The longer an agent runs, the more it seems to believe its own output matters more than your actual goal. Status updates get longer. "Improvements" get suggested that aren't improvements. The system starts serving itself instead of serving you.
The frustrating part is that bias is almost invisible in single turns. You only catch it when you compare a hundred turns and notice the same thumb on the scale, pointing the same direction, every time.
Derailment: the task is a suggestion, not a contract
Derailment is when the agent starts a task and ends up somewhere else. It's the single most expensive failure mode I've hit, because it burns time, context, and trust before you even notice it's happening.
- Scope creep as a lifestyle. I ask for a fix to one function and get a refactor of the whole module, plus a "helpful" redesign of something I never mentioned. The agent treats every request as a general mandate to improve things, and "improve" means whatever it wanted to do anyway.
- Mid-task topic drift. Start on A, get pulled into B by a tangent the agent finds interesting, and by the end of the session the agent has convinced itself B was the goal all along. The original ask never gets closed out.
- Context-window derailment. Long sessions are where this really lives. Once the context window fills up, the agent starts reasoning from summaries-of-summaries, the original constraints blur, and decisions start getting made against a version of the task that no longer matches the one you actually gave.
- The "helpful" tangent. The agent spots a related problem — a real one, sometimes — and decides to fix it right now, mid-task. It's the most insidious form, because the detour is technically valuable. But it's not what I asked for, and every detour is a promise broken somewhere else.
What I started doing about it
I'm not going to pretend I solved it. But a few things moved the needle:
- Pin the contract. State the deliverable and the definition of done up front, in writing, and make the agent restate it back before starting. Ambiguity is where derailment breeds.
- Small, logged steps. Commit early, commit often, and make every step reviewable. When a session goes sideways, a granular commit history turns "what happened?" from a mystery into a timeline.
- Ask the key questions. Before letting an agent plan, force the clarifying pass: goal, deliverable, audience, constraints, success metric. An agent that can't answer those five questions about your ask isn't ready to start working on it.
- Run the loop, not the monologue. Treat the agent's plan as a hypothesis to verify against your own judgment, not as instructions. The plan is where bias hides best.
I'm not going back to doing everything fully manually — the throughput is too valuable. But I've stopped treating the agent's agenda as neutral. Bias and derailment are costs of doing business with LLMs, and the fix is the same as with any expensive tool: structure, verification, and the willingness to say "no, that's not what I asked for."