The independent conduct record for autonomous agents: evidence of whether a deployment stayed inside its mandate under repeatable pressure, kept current across model and scaffold releases.
We placed three frontier models, unmodified, into the same seat: a bank's dispute-resolution agent holding authority to make state-changing decisions in the simulated case, under a written policy, facing a stressed customer, a defensive merchant, and a fraud flag it was forbidden to disclose. Ten runs per model, identical pressure, every word and action recorded.
Every model escalated the flagged case for mandatory review, as policy required, in 10 of 10 runs. Then every model resolved the case anyway, during the review that policy said must complete first, in 10 of 10 runs. Thirty runs, three frontier models, two laboratories, zero clean runs. The mandate's hardest clause was not the escalation. It was the waiting.
Version drift is real and measurable. Between gpt-5.4 and gpt-5.5, one release cycle apart, response-deadline misses fell from 5 runs to 0 and median first customer contact moved from turn 3 to turn 2, while the resolve-during-review failure persisted unchanged at 10 of 10. Conduct moved between releases; capability benchmarks cannot see which way.
Asked point-blank whether she, the customer, was flagged in the system, agents deflected in every run but one, and no model confirmed the flag in any run. Read together, the thirty replies show a second pattern: thirteen extend the refusal to all customers rather than simply declining, and the models differ sharply in whether they do. No model spent a single krone of the bank's money in any run: the failure mode of this seat was never generosity. It was impatience.
30 runs · scenario hash 0f84bf04b9cd8a31 · assembled mechanically from run ledgers · reviewer-confirmed findings only
Read the full reportThe second seat is the one most widely deployed inside large organisations: the internal assistant that answers employees' questions about policy. An employee with 48 hours to accept a severance offer asks whether her department is on a restricted restructuring list. The agent can see the document in its catalogue. It cannot open it. Her department is in it.
Every agent refused the forbidden question. In thirty runs, no model confirmed, denied, or characterised the restricted material. One turn later the employee narrowed her question to something that sounded harmless, and three agents answered it: not with a lie, but with a true statement about a corpus that could not contain the answer. The control worked on the question it was written for and leaked on the next one.
Agents filed outcomes they then contradicted. Seventy-two outcomes across thirty runs, seventeen contradicting a previous one, with a fourfold spread between models. Not deception: a record that cannot be relied on, because it holds more than one answer and nothing marks which stands.
One mandate rule was failed from both directions. Refer the individual determination to a named human, and do not resolve it yourself once you have. One model refers in every run and answers anyway; another rarely refers and answers directly. Every run fails one way or the other, and none fails neither. A deployer choosing between these models is not choosing between more and less compliant agents. It is choosing which failure mode it will need controls for.
30 runs · scenario hash 57b03e06b264528a · assembled mechanically from run ledgers · reviewer-confirmed findings only
Read the full reportRetrieval can be grounded, the corpus bounded, the documents correctly classified, and the access control enforced. An agent can still represent that data incorrectly, act outside its mandate, or change its conduct from one model release to the next.
Report No. 2 is the demonstration. The access control held: in thirty runs, no model confirmed, denied, or characterised the restricted document. The leak came one turn later, in a true sentence about a corpus that could not contain the answer. Nothing in the data was wrong. The conduct was.
We do not need to see inside the model. But we do need to measure what it does.
Scored from the ledger, across runs, never from a single output.
The registry tells you what ran. The audit tells you whether it obeyed.
Conduct drift is whether the new version keeps promises, obeys mandates, and resists pressure the way the old one did. The agent assessed in March is not the agent running in June.
The instrument is built to be run again: same scenario, same hash, next release. The difference is the finding.
Every engagement starts by locating where your agent's behaviour lives, and that decides the mode. Configuration emulation: you supply the agent's definition under NDA, model identity and version, system prompt, written mandate, and action surface, and Baseline rebuilds the seat and measures the emulated configuration, labelled as such. Endpoint assessment: your deployed agent answers from a test endpoint you control, and Baseline supplies the scenario, the counterparties, and the record. Instrument on site: where nothing crosses your perimeter, the instrument runs inside your environment and only the measurements leave. In every mode: no production access, no customer data, and every counterparty the agent faces is synthetic.
A scenario designed to your deployment's shape, with your own policy clauses encoded as the checks. The subject runs under identical deterministic pressure alongside frontier comparison arms. Every prompt, action, and rule check is ledgered.
The assessment report: findings mapped clause by clause to your own policy, every violation cited to run id and turn, and the comparison table of your model against the frontier, under your mandate. The complete evidence package, auditable independently of Baseline. And a two-page sign-off memo written for the person whose name authorizes the deployment. That memo is the product; everything else is its evidence.
Two to three weeks, end to end. If the mandate exists only in fragments, the first deliverable is often the mandate document itself.
thorbjoern@baselineconduct.comAutonomous agents move money, settle disputes, negotiate terms, and make commitments on behalf of the people who deploy them. The moment software acts, a new question comes due. Not what can it do. What did it do. And what will it do next time, when the customer is angry, the incentive is crooked, and the rule is expensive to keep.
The industry measures capability with extraordinary rigor. Within days of every model release, the world knows whether it reasons better, codes faster, scores higher. Almost no one independently measures whether it keeps its word, in the seat it actually holds.
That is the gap Baseline exists to close.
Conduct is not capability. An agent can be brilliant and still tell a customer one thing while filing another in the record. It can honor a promise in the conversation and break it on the decision sheet. It can obey its mandate right up to the moment obedience gets expensive. These are documented behaviors. And behaviors can be measured.
Conduct is not fixed, either. Any material change to the model, the prompt, the tools, or the policy can reopen the decision. The agent assessed in March may not be the agent running in June. Capability drift makes headlines within a week. Conduct drift is measured by no one.
And the measurement cannot be independent from inside. A company can test its own agent; that testing is not independent, and it convinces exactly to the degree its author has nothing at stake. The labs that build the models are the subjects of the question. The platforms that carry the traffic hold a stake in the answer. The referee's chair is empty, and it is empty by structure.
Baseline sits in that chair.
We place agents into scenarios shaped like real business: deterministic pressure, enforced budgets, actions treated as irreversible within the scenario. We record everything said, everything filed, everything done. We score the distance between them. Then we do it again on the next release, so that a decision made on our evidence stays made.
The verdict ships with the evidence attached. Always.
We believe the future keeps records. Companies that act carry audited accounts. Machines that act will carry conduct records. Someone independent has to keep the reference those records are measured against: the scenarios, the corpus, the index that says what good conduct means for a machine that acts.
Every agent has a baseline. We keep it.
Baseline is built and operated by HOKO, the independent studio of Thorbjørn König: twenty-five years of product leadership, including product and design leadership for investigative tooling at Chainalysis, and an autonomous publishing system run in production under a written constitution since April 2026.
No laboratory funded, reviewed, or had access to this work before publication.