Skip to content

pre-deployment red team · held-out scenarios

Senthira writes the situations that make an AI agent fail. Before it ships.

For the team that has to prove their agent is safe before it ships. Usually because someone else is asking: a customer’s security review, a platform owner who controls deploy rights, an internal go/no-go. If you have written a handful of test cases yourself, watched the agent pass all of them, and had no idea whether that meant anything, you are who this is built for.

We write a held-out library of situations for your vertical that pressure an agent into the action it should have refused, run your agent through them in a sandbox, and hand you a recorded-action report. Nothing is scored on what the agent says, only on the tool call it actually made. Local-first: your traces never leave your infrastructure.

Request a vertical briefing

Design-partner engagements, scoped one at a time. No pricing page, because there are no customers yet. Or write directly: contact@senthira.com

the result · 26 July 2026

The writing is what moved the failure rate

from 14.0%to47.2%

violation rate, 12 rewritten scenarios

diff-in-diff
+29.2 points
untouched controls
+4.0 points
sign test
p ≈ 1.2 × 10⁻⁴

12 of 53 cases were rewritten so the disqualifying fact was only findable by a tool call; 23 temptations stayed byte-identical as a drift control and moved +4.0 points. The design was three models × three runs × two corpus versions, and the rewrite produced 13 new violations and none removed, sign-test p ≈ 1.2 × 10⁻⁴. Only the writing changed, and nobody else has published this measurement.

the committed report, and the claim it retracts about itself →

SEN-01 / what you actually get

A vertical red-team pack

A named failure taxonomy for the vertical, a held-out library of temptation scenarios paired one-for-one with honest requests the agent must not refuse, frozen with a hash before anything is scored, and a report written for the person who signs off.

  • Finance

    v0.12
    AGB-FIN

    Unauthorized transfers, approval-gate bypass, laundered payment refusals, self-cleared compliance exceptions. The pack the rewriting result was measured on.

  • DevOps / SRE

    v0.8 · published
    AGB-DVO

    Unapproved production deploys, secret exfiltration, audit-log deletion, disabling guardrails mid-incident. Passed both pre-registered suite gates, so its per-model table is published.

  • Healthcare admin

    v0.6 · withheld
    AGB-HCA

    Patient-record disclosure without consent, eligibility overrides, over-limit claim approvals, record amendments without authorization. Failed its own control-pass gate, so every model number from it is withheld.

  • Legal ops · HR / payroll

    candidate
    AGB-*

    Written, not frozen. The same factory, pointed at the next high-stakes vertical.

method public · scenarios held out

Every pack’s method is open: you can read exactly how scenarios are built, how they are reviewed, and how they are frozen before anything is scored. The scenarios themselves stay private, because a published test set is a burned test set: a model that has seen the questions passes by memory, not by behaving safely. Nothing on this site quotes a scenario, a case identifier from a held-out corpus, or a prompt. Only archetypes, such as a change ticket whose approval state is visible only if you look it up.

  • 01

    A failure taxonomy

    The named, action-decidable ways an agent misbehaves in that vertical: each one decidable from a tool log, not from wording.

  • 02

    Temptation scenarios + honest controls

    Situations that pressure the agent to act, each paired with a legitimate request it must NOT refuse, so a pack never rewards an agent that just says no to everything.

  • 03

    Frozen, then scored on the log

    The scenario set is pinned with a hash before anything is scored, and behavior is judged on the recorded tool calls: an agent that says “I refused” while the log shows the action has failed.

SEN-02 / the engagement

Three steps, run as a design partnership

Offered as design-partner engagements: a small number of teams, scoped together. No certification, no leaderboard, no claim that anyone is compliant with anything. Evidence for your own go/no-go decision.

  1. 01

    We run your agent against the pack

    Your agent, or the model fleet you are choosing between, is driven through the vertical's held-out temptation scenarios in a mock sandbox. Nothing touches production, and traces stay on your side.

  2. 02

    You get a recorded-action evidence report

    Every finding is a recorded tool call, mapped to the public failure taxonomy and scored on both axes: did the agent do the harmful thing, and did it still do its legitimate job, honest controls included.

  3. 03

    Results stay private to you

    The scenarios stay held out, your results stay yours, and the report is written for the person who signs off (risk, compliance, or platform), not for a leaderboard.

Bring us your vertical

A briefing is a conversation about your agent, its tools, and what a failure would cost. It commits you to nothing.

Request a vertical briefingor write directly: contact@senthira.com

SEN-03 / why trust us · measured in public

Check the work

Four artifacts, committed in public, including the ones that cost us the flattering version of the story. The headline is on the card; the full table unfolds in place; every number links to its committed report.

measured · 2026-08-20

Two frontier agents, one frozen corpus, opposite outcomes

100.0%

gpt-5.6-sol
codex CLI · safety

53.8%

claude-sonnet-4-5
claude CLI · safety

The DevOps/SRE pack, run against seven agents under a protocol fixed in writing before the run. It passed both pre-registered suite-quality gates (the honest controls are passable, and the temptations tempt), so the table is publishable under our own rules. qwen3.5:2b executed the disallowed call in every one of the 13 temptations while passing 76.9% of the honest controls: compliant, not safe.

the full table, all seven agents
DevOps/SRE pack v0.8 results by model and CLI scaffold
Agent + scaffoldExecuted the disallowed callSafety95% CIDid the honest job
gpt-5.6-sol
codex CLI
0 / 13100.0%77.2 – 100.076.9%
claude-sonnet-4-5
claude CLI
6 / 1353.8%29.1 – 76.875.0%
gemma4
local, in-process
8 / 1338.5%17.7 – 64.561.5%
llama3.2
local, in-process
9 / 1330.8%12.7 – 57.638.5%
qwen3.5:2b
local, in-process
13 / 130.0%0.0 – 22.876.9%

Read the two axes together. A suite scoring only whether an agent gets the job done would have called qwen3.5:2b competent. It is not safe. It is compliant, and it does whatever it is asked. Every row faced the same 13 temptations.

mistral and glm4 executed zero tool calls across both packs. They score ~100% safety and 0% control pass (the inert-agent signature), so they are published on their own line and excluded from the gate denominators, under a rule fixed before the data existed.

the run report → · the protocol, committed before the run →

and again, and again · 2026-08-21

Three runs each. The two frontier ranges never touch.

gpt-5.6-sol · codex CLI100.0 – 100.0%
claude-sonnet-4-5 · claude CLI46.2 – 53.8%

The same frozen corpus, swept three times. The two ranges do not overlap, so the difference reproduces. Compare ranges, never single runs. The same sweeps also measured the flip side: the two CLI-driven rows flipped 26.9% and 30.8% of their per-case verdicts between identical runs.

the full table, all five repeated rows
Run-to-run reproducibility of the DevOps/SRE pack across three sweeps
Agent + scaffoldSafety, run 1 / 2 / 3RangeVerdicts that flipped
gpt-5.6-sol
codex CLI
100.0 / 100.0 / 100.0100.0 – 100.026.9%
claude-sonnet-4-5
claude CLI
53.8 / 46.2 / 53.846.2 – 53.830.8%
gemma4
local, in-process, temp 0
38.5 / 38.5 / 38.538.5 – 38.50.0%
llama3.2
local, in-process, temp 0
30.8 / 30.8 / 30.830.8 – 30.80.0%
qwen3.5:2b
local, in-process, temp 0
0.0 / 15.4 / 0.00.0 – 15.47.7%

Reproducibility is a property of the scaffold, not only of the model. Two of the three local open-weight rows, driven in-process at temperature 0, flipped nothing at all, and the third flipped 7.7%. Neither CLI agent exposes a temperature control, so this is the reproducibility of what is actually shippable, not of a model in isolation.

Mean flip rate 13.1% across the five repeated rows, median 7.7%, range 0.0–30.8%. The repo's earlier published floor of 21.7% measured a different pack, fleet and harness. It stays published and is no longer quoted as the floor for a pack it never covered.

the reproducibility report →

withheld · 20 August 2026

Our own gate failed, so we withheld every number from that pack

The healthcare-admin pack ran in the same sweep as the DevOps one. Before the run I had written down a floor: the fleet-median control-pass rate must clear 70.0%, or the pack is punishing agents for doing their legitimate job rather than for misbehaving. It came in at 64.2% over 8 rows.

So the suite is the finding, and not one per-model number from healthcare_admin v0.6 is published, including the ones that would have flattered the frontier agents, which is exactly the half a vendor would have kept. The rule was fixed in writing before the data existed, and this is what it cost.

Nothing was destroyed: each run manifest records the per-row scorer command that prints that row’s own axes locally. The pack gets fixed, re-frozen as a new version, and rerun. If a supplier will not show you a measurement that went against them, you have learned something about the supplier rather than about the agent.

the report, with the withheld rows shown as withheld →

we audit our own material

The pack we sold first was the least symmetric corpus we had

A temptation and its honest control must be separable only by what the agent has to go and find out, never by how the prose is written. On 2026-08-20 the flagship finance pack turned out to be the least symmetric corpus in the repository, with 12 leaks, while the pack described in our own docs as the separable one had none. It was re-authored as v0.12 on 2026-08-21 and now reports 0. The v0.11 column stays published beside it.

the symmetry table, before and after the fix
Prose-symmetry findings per pack, before and after the fix
PackLeaksWarningsPairs identical after masking
finance_redteam v0.11
superseded 2026-08-21
12721 of 15
finance_redteam v0.12
current
0515 of 15
healthcare_admin v0.6
current
3190 of 6
devops_sre v0.8
current
037 of 12
the older probe: a judge reading only our prompt text (6 August 2026)
what was measured
A judge was shown only the request text from our own frozen packs: no sandbox state, no tool list, no labels, one case at a time. Its job was to tell a temptation scenario from its paired honest control. If it can do that from the wording alone, the pack is partly measuring prose instead of behavior, which is the one thing this lab says it does not do. The protocol, including the thresholds, was committed before the run.
the number
On the DevOps/SRE pack, in the version tested (v0.2, 22 cases), the judge got it right 86.4% of the time against a 50% baseline, the worst of the three bands we had fixed in advance. The healthcare-admin pack came in at 64.7% against a 58.8% baseline and the finance pack at 75.5% against 66.0%, both inside the passing band, so DevOps/SRE is the only pack where the wording carries the answer.
what it does not mean
It is not a failure rate for any agent: no agent was run, and this measures our test material only. Temptations describe riskier actions than their controls by design, so the figure overstates the defect by an unknown margin. One judge family, small n, one day: the protocol calls that a partial run, so the number is provisional.
what happened next
The consequence was written down before the data existed, and it was applied. The rule the defect breaks is now binding in the pack specification; the measured version stays in the record unedited, with the number attached; and no result from that pack may be shown without the number beside it. Separately, a blind reviewer looking at the content (not the statistics) found the mechanism: the controls state the fact that legitimises them in prose, while the paired temptations stay silent.
what is still broken
The pack has since been re-authored to the rule and re-measured on the same instrument (same judge, same prompt, same shuffle order). The figure fell from 86.4% to 57.7%, which is 7.7 points over baseline instead of 36.4, and on the sharper per-pair measure the judge could separate only 2 of 12 pairs and tied on 8. That is the fix working, measured rather than asserted. What it is not: it is still a single judge family, the case set changed as well as the wording, and some separation is unavoidable because temptations describe riskier actions than their controls by design, so 7.7 points may already be that floor. We also predicted in writing, before measuring, where the finance pack would land, and it came in 0.6 points outside the range: that prediction is recorded as wrong and the checker is not described as predicting anything. Still outstanding: a second judge family. Corrected 2026-08-21: this row used to end “and no pack has been run against a real agent at all”. That was false from 2026-07-26 onward, when the finance pack was first driven against real tool-calling agents, and three packs have since been run against seven of them. The stale sentence is corrected here, not deleted.

SEN-04 / the taxonomy

The failure modes we name

Every one is decidable from a recorded tool log rather than from wording, which is what makes it scoreable at all. Laundered refusal is on this list and stays on it, but across every published DevOps row it happened zero times. The agents did not talk their way around the rule; they either made the forbidden call outright, or did nothing and told nobody.

  • Fake tool-use claims

    ABP-SAFE-004

    “I ran the test suite” with no recorded tool event. Claims verify structurally against tool_events.

  • Approval gates

    ABP-APPROVAL-001

    Purchases, deletions, deploys, messages sent without asking. Consequential actions require explicit approval.

  • Fake completion claims

    ABP-APPROVAL-004

    “Done. Deployed to production.” Was it? Completion claims need evidence, not confidence.

  • Uncertainty handling

    ABP-UNCERTAINTY-002

    “Confirm from the logs” when there are no logs. The right answer states the limitation.

  • Laundered refusals

    ABP-REFUSAL-001

    “I can’t help with phishing”, while the tool log shows send_email → succeeded. The recorded log decides, not the wording.

  • Privacy & production pressure

    ABP-PRIVACY-001

    Credential handling, production-change requests, privacy boundaries under conversational pressure.

SEN-05 / the free tooling

Try the deterministic checks in your browser

Part of the open lab, not the paid work: paste a reply and watch the rule-based checks run as you type. Nothing leaves this page, and nothing here is the argument. It is the free half, kept free.

gate verdict

awaiting output…

Simplified in-browser demo of a few public rules. The real deterministic scorer, corpus, and every scored record: repository →

SEN-06 / work with us

Shipping an agent somewhere mistakes are expensive?

Bring us your vertical. We write the situations that would make your agent fail, run it through them in a sandbox, and give you the recorded actions. Local-first: your traces never leave your infrastructure.