Evidence first: designing an AI agent that operations teams can trust
Design rules from an agent that finds the root cause of enterprise network problems, and why each claim must point to its evidence.
A network problem rarely arrives as a clear fault. It arrives as a complaint: “the Wi-Fi is slow on the third floor”, or “I cannot log in since this morning”. An engineer then checks the controller, the access points, the authentication server and the recent changes, one tool at a time.
A language model can do this search faster. It can also produce a confident answer with nothing behind it. For an operations team, that is worse than no answer, because they must act on it.
Bullseye Zero is an agent that I am building for this problem. These are the design rules that make its answers something an engineer can check.
Missing evidence is unknown, not healthy
The most dangerous error in monitoring is silence that looks like health. A data source times out, a query returns nothing, a score is missing, and the dashboard shows green.
In Bullseye Zero, every part of the network starts as unknown. A part becomes healthy only when evidence says so. A failed or empty query produces an unknown status with a reason, and the final report names every dimension that it could not check.
The model chooses; the code decides
The language model picks the next step: which check to run, which question to ask, or when to conclude. It does not compute the answer. Deterministic code sets the state of each device, ranks the possible causes and computes the confidence.
This split keeps the reasoning flexible and the conclusions reproducible. If the model picks a step that is not allowed, or tries to conclude below the confidence bar, the step does not run, and the record says why.
Every claim points to a record
Each piece of evidence is stored with the same fields: the source, the exact query, the time window that was requested, the time window that the source actually returned, the status, and a hash of the raw data. The raw data is kept.
The report is designed to link each claim to the record behind it, and to mark any claim without one as uncited. An engineer who doubts a conclusion can then open the evidence and see exactly what the system saw. An auditor can see it a year later.
A measurement is true only from where it was made
An access point reports its view of the radio. The authentication server reports its view of the login. Neither sees everything. The agent keeps every result with its vantage point, and it does not merge them into one number.
When independent vantage points agree, confidence rises. When they disagree, the disagreement is itself evidence.
Use the structure of the network to rule out causes
An access point on another floor cannot explain one user’s slow session, however bad its own health score is. The agent builds a dependency map of the network and considers only faults that lie on the path of the symptom.
This rule needs no historical data, and it removes most wrong answers before any ranking starts.
Observe, do not change
Every connection to the network is read-only. The agent investigates and recommends; people make changes. For a system that operates inside a customer’s network, this is not a limitation. It is the condition for being allowed in at all.
Let the agent ask a person
Some facts are not in any system: “Did this start after the office move?” The agent can stop, ask one question, and wait. The answer is stored as evidence like any other record, and the case resumes when it arrives, even after a restart.
What this means for enterprise AI
These rules make the agent slower to conclude than a model that simply answers. They also make every conclusion checkable. In an enterprise, that trade is almost always correct: a system that people can audit is a system that they will actually use.