Matthew Andrews
Technology & Research

Can a System Learn Not to Make the Same Mistake Twice?

By Matthew Andrews ·

Concepts for learning from mistakes in multi-agent workflows

One of the questions driving my work is simple: if a system can recognize that it made a mistake, can it use that understanding to avoid repeating it?

I’ve encountered versions of this problem in programming, technical leadership, and the workflows I’ve built for myself. A problem gets identified. Someone corrects it. Work continues. Later, a similar problem appears, and someone has to intervene again.

The immediate issue may have been resolved, but the process hasn’t necessarily improved.

That gap is what interests me. It is also the focus of my capstone and planned research into multi-agent architectures: how systems might diagnose mistakes, learn from them, and improve their workflows with less human intervention.

Fixing an Output Is Only the Beginning

When a system produces an incorrect result, correcting that result is useful. Understanding why it happened is a different task.

Did the system misunderstand the request? Did it lack necessary information? Did one part of the workflow pass incomplete information to another? Was there a check that should have caught the issue earlier?

Those possibilities matter because they call for different corrections.

A system that lacks information may need to ask a question. A system that misinterprets information may need a better decision process. A system that fails to check its work may need a verification step.

If every failure leads to the same response—try again—the workflow can spend more time producing results without becoming more dependable.

I want to explore what happens when diagnosis becomes part of the workflow itself.

Why Multiple Agents?

My research centers on multi-agent architectures, where different agents can take responsibility for different parts of a task.

One possible arrangement is to separate producing a result, evaluating it, and diagnosing a failure. That separation creates an opportunity for a result to be examined from more than one perspective.

It also introduces new problems.

Agents can misunderstand each other. They can pass along incomplete assumptions. Several agents can agree on an answer and still be wrong. Adding more participants can increase the amount of work without improving the outcome.

For me, that makes coordination part of the research question. What should each agent be responsible for? What information should move between them? When should a workflow continue, retry, or stop?

The architecture needs to earn its complexity through better results.

Remembering a Mistake Isn’t Enough

A system could store a record of every error it encounters and still repeat those errors.

The harder question is whether it can turn an experience into a useful lesson—and recognize when that lesson applies.

Imagine a workflow that fails because an input is missing. It might record a lesson to check for that input before proceeding. That could be helpful in similar situations.

But what if the input is optional in another context? Applying the lesson too broadly could create a new mistake.

Learning therefore requires judgment about scope. A correction needs to address the original problem without becoming an instruction that distorts unrelated work.

This is one of the aspects I find most interesting: determining what a system should retain, what it should revise, and what it should avoid generalizing.

Improvement Has to Be Demonstrated

A system describing its mistake convincingly doesn’t prove that it understands the cause. Changing its workflow doesn’t prove that the change helped.

The standard I want to work toward is practical: does it perform better when it encounters the problem again, including when the details are different?

That means examining more than whether a second attempt succeeds. I want to know whether the correction carries into a new situation, whether it creates problems elsewhere, and whether the improvement justifies the additional time and resources.

It also means being willing to reject a proposed correction.

A self-improving system needs a way to distinguish useful changes from changes that merely sound reasonable. Otherwise, it risks accumulating instructions and complexity without becoming more reliable.

Where People Still Matter

My longer-term goal is to develop workflows that can diagnose and address recurring problems without requiring a person to intervene every time.

Reaching that goal requires clear boundaries.

A system needs to recognize when it lacks enough information to continue. It needs a way to surface uncertainty and make its decisions understandable. Some situations will still require a person’s judgment, particularly when an error could have significant consequences.

Those questions belong in the design from the beginning. They affect what the system is allowed to change, how a correction is evaluated, and when human review is necessary.

Reducing repetitive intervention is valuable only if the resulting workflow remains dependable.

What I’m Working Toward

My software engineering studies at St. Mary’s University have given me a place to explore these questions through agentic systems and workflow optimization. That work has also helped shape the foundations of OmniMint.

The research is ongoing. I’m describing an objective I’m working toward, with questions that still need to be answered through implementation and testing.

What draws me to this work is the possibility of making improvement part of a system’s normal operation: identifying a problem, examining its cause, testing a correction, and carrying forward what proves useful.

I want to understand how far that process can go—and what it takes to make it trustworthy.

Research Overview

Status: Ongoing Capstone & Planned Research
This is a description of the research direction, not a report of validated results or a published paper.

The Question

How can a multi-agent workflow use its diagnosis of a mistake to improve future work, without turning a narrowly useful correction into a rule that causes problems elsewhere?

A Workflow to Investigate

  1. ProduceGenerate a candidate result for a defined task.
  2. EvaluateCheck the result against the task’s requirements.
  3. DiagnoseInvestigate the cause when a check fails.
  4. Test a CorrectionTry a scoped change and evaluate its effects.
  5. Retain What HelpsCarry forward corrections supported by evidence.

This conceptual sequence illustrates the questions I’m exploring; it does not claim that each capability has been implemented or validated.

What Would Count as Improvement?

The evaluation should examine whether a correction helps on a new version of the problem, whether unrelated tasks still work, and how much extra time and computation the workflow requires. Comparisons with the unchanged workflow would help distinguish improvement from a successful retry.

Limits and Open Questions

Several agents can share the same wrong assumption. A convincing diagnosis can be mistaken, and a correction can apply too broadly. Human review remains important when uncertainty or consequences exceed the workflow’s boundaries.

Evidence and Publication Status

No experimental results, dataset, public demonstration, or peer-reviewed publication is presented here. Methods and findings will be added when they are ready to share and can be supported with evidence.

Back to Blog