Welcome to Part III of my article series where I take you along on the journey to building an AI personal decision system agent called “What Do I Do?”.

In my previous article, I introduced the concept of the decision state machine which maintains the user’s thoughts as a state object in the decision making process. The state object is meant to act as the agent’s working memory. After each turn, it updates what it knows about the user’s goal, constraints, values, options and confidence. For example,
{
“goal”: “…”,
“constraints”: [],
“values”: [],
“options”: [],
“criteria”: [],
“confidence”: 0.7
}
In theory, this gives the agent continuity. In practice, my first implementation exposed just how fragile that continuity can be.
My original plan was for this article to cover the evaluation system. But implementing the decision state machine was far less straightforward than I expected, and I realized I needed to understand the agent’s failure modes before building the full evaluation layer. This also aligns with a point Hamel Husain has made in his writing on LLM evaluation, before building an eval system, you should carry out some error analysis. It helps to first understand how and where the system is actually failing.
I’ll tell you more about that in this article.
Learnings from Implementing the Decision State Machine
To not miss out on any new articles, consider subscribing.
Following the previous article, I began implementing the decision state machine and while testing it out with sample questions, the agent’s performance was much worse than I expected.
Here are some of the ways the agent struggled:
- State failures: missed values, lost weather information, contradictions between assumptions and uncertainties.
- Conversation-policy failures: repeated questions, too many clarifying questions, backtracking.
- Output-quality failures: generic/passive responses, overly verbose recommendations, parroting the user’s answer without adding value.
- Infrastructure failures: rate limits, timeouts, malformed JSON, schema-validation failures.
State failures:
- Not extracting options, values, and fears correctly: The agent struggled to extract values for the decision state machine. In this example, we see the agent miss the user’s response stating their preference for a formal business suit. The agent still asks the user to pick between a formal business suit and another (opposite) option again.
Similarly, here, I had mentioned weather as my main goal when choosing what to wear to work the next day. However, in the rest of the conversion down to the final recommendation, the agent does not mention or consider weather as a factor at all. That information just got lost.


- Contradictions: This showed up in extracting options, clarifying questions, and in the final recommendation as well. In this example below, the agent had contradictory beliefs about the user’s assumptions and uncertainties. We see the agent state an assumption of preferring sweet flavors over savory ones (extracted from user’s response), but still list the flavor preference (sweet vs. savory) as an uncertainty.

Conversation-policy failures:
- Repeating clarifying questions: This is self-explanatory and the most basic failure the agent displayed. It sometimes simply asked exactly the same clarifying questions in a row.

- Asking too many clarifying questions: The agent struggled to get to a recommendation, even for a simple decision, and kept asking clarifying questions, sometimes even up to 8 clarifying questions before it finally gave a recommendation. Even as the creator of the agent, I was getting pissed off and frustrated, let alone a user.

To not miss out on any new articles, consider subscribing.
Output-quality failures:
- Verbose recommendation: As a user, if an agent gave me this long paragraph when I needed help in just deciding what to wear to work tomorrow, I would probably just zone out and close the tab. The agent needs to go straight to point and give a final recommendation that the user can easily read, understand, and take action on.

- The responses were too generic and passive: In this example, the first response from the agent was really generic. It was not personable and lacked warmth, which could definitely be improved.

- Repeating the user’s response as a recommendation verbatim with no extra information: Following from the previous screenshot (same decision), not only did the agent backtrack/contradict itself down the line, but it also repeated my preference for “keepsake” as the final recommendation, with no added information, context, or examples that add value to me as a user.

Infrastructure failures
- Model and runtime failures: Sometimes, the model simply failed. Due to model rate limit, timeout errors and retries, json validation failures, etc.

What These Failures Revealed
One takeaway for me in this process is that a working LLM call is not the same thing as a working agent. Once an LLM has to maintain state across turns, decide whether it knows enough, infer implicit preferences, recover from model failures, and produce an appropriately-sized answer, reliability becomes a systems problem.
To not miss out on any new articles, consider subscribing.
How I Mitigated These Failures
- Built a structured LangGraph decision workflow: I broke the decision process into a structured LangGraph workflow: extraction → clarification → option generation → evaluation → stress test → recommendation. Instead of relying on one general-purpose model call to handle everything, each stage now has a narrower responsibility and clearer transition criteria, making the agent more explainable.
- Added stake-aware clarification limits: For now, I use stake-aware question caps as a simple stopping policy. Longer term, I want the agent to estimate whether another question is likely to materially change the recommendation.
- Made clarification prompts warmer and more conversational: I rewrote the clarification prompts so they feel less like a questionnaire and more like a natural conversation. The goal was to make follow-up questions feel context-aware, concise, and proportionate to the decision being made.
- Simplified extraction and recommendation schemas: I reduced the number of fields the model had to populate and made each field more clearly defined. This lowered the cognitive load on the model and reduced cases where it produced inconsistent, redundant, or poorly structured state.
- Limited extraction context to recent messages and relevant state: Instead of repeatedly sending the entire conversation history, I now provide only the most recent user messages and the parts of the existing decision state that are relevant to the current turn.
- Added JSON-object fallback when strict schema validation fails: When the model fails to produce output that matches the expected schema exactly, the system can fall back to a more flexible JSON response instead of failing the entire step. This gives the workflow a chance to recover gracefully from formatting or validation errors.
- Added model fallback routing and token budgets for the different steps in the workflow: I introduced fallback models for cases where the primary model hits a rate limit, times out, or fails unexpectedly. I also set different token budgets for extraction, clarification, and recommendation steps so simpler tasks do not consume unnecessary context or output tokens.
- Implemented better error handling and diagnostics logging for debugging: I added explicit exception handling around model calls, parsing, and validation, along with more detailed logs for each step in the workflow. This made it much easier to distinguish between model-quality issues, schema failures, orchestration bugs, and infrastructure problems.
These changes don’t prove that the agent is now reliable. However, they give me a better architecture to test. The next step is measuring whether each intervention actually reduces the failures I observed.
To not miss out on any new articles, consider subscribing.
Conclusion
While experiencing these bugs was frustrating, they became a learning point for me. The bigger lesson was that getting an agent to work reliably was not just about getting the right prompt. It also meant careful state management, orchestration, stopping rules, schema design, fallback strategies, and observability.
More importantly, these failures have now given me something concrete to evaluate and have formed a base for the eval metric set for the agent. Repeated questions, lost state, contradictions, excessive clarification, verbosity, and structured-output failures can all become measurable behaviors in the evaluation system. So instead of asking only, “Does the agent work?”, the next question becomes: “How reliably does it work, and where does it still fail?”.
That’s what I will cover in the next article.
Thank you for reading.
Aniekan
To not miss out on any new articles, consider subscribing.
