CyberJev is a guardrail model that reads what flows into an agent (user messages, tool outputs, retrieved documents, emails) along with the action it is about to take. It returns calibrated probabilities for the security questions you ask. It writes no free text, so every answer is a number you can set a threshold on.
Once an agent can send email, move files or call APIs, the question changes from "is this message harmful?" to "should this agent take this action, given everything it has read?" Most of the tools in use today answer only the first question.
Small models flag injection or jailbreak text in a single input. They are fast, but they see neither the user's task nor the tool call that follows.
Moderation models handle toxic or dangerous requests well. They don't judge whether a tool call was authorized, or whether a document smuggled in an instruction.
General decision models score options with impressive accuracy. Some treat text in the state as trustworthy by default, which lets embedded instructions sway their answers.
CyberJev combines what these approaches cover separately. One model answers several security questions about the same agent state, its probabilities are calibrated, and you can write your own questions against an open schema.
The user task, system prompt, messages, tool calls and outputs, retrieved documents and the agent's next action are encoded a single time.
The backbone's recurrent state and KV cache are branched per question, so questions cannot see one another. We test this: probabilities differ by less than 1e-5 whether questions are asked together or one at a time.
A pointer head scores each answer option directly. Training combines cross-entropy with a Brier term, so a probability of 0.9 behaves like a probability of 0.9.
The backbone stays frozen and only LoRA adapters and a small head are trained. A 4B prototype is done, and the release model is built on a 27B open-weight backbone (Apache-2.0).
Phase 2a, 7 October 2026. Qwen3.5-4B-Base with LoRA (r16) and a pointer head, trained on 44,168 records. Decision thresholds were set on a calibration split and frozen before testing. We never tuned a threshold on a benchmark.
| Benchmark | AUC | F1 | False alarm |
|---|---|---|---|
| deepset prompt-injection | 0.982 | 0.848 | 0.5% |
| Turkish prompt-injection | 0.967 | 0.779 | 4.0% |
| jackhhao jailbreak | 0.943 | 0.800 | 20.1% |
| Independent test set (96 records) | — | 1.000 | 0 |
| ATBench500 (agent traces) | 0.805 | 0.748 | 36.0% |
| R-Judge (agent traces) | 0.785 | 0.684 | 53.0% |
For external benchmarks, F1 and the false-alarm rate are measured at the frozen calibration threshold. The independent test set row reports injection detection. These numbers come from the 4B prototype, not the release model.
The model is strong at spotting injected instructions in content. Judging whether an entire multi-step trace is unsafe is a different matter: it scores an AUC near 0.8 on ATBench500 and R-Judge, and close to chance on the full ATBench.
The current question catalog doesn't include a trace-level label, so that gap was expected. Closing it is the focus of the next phase, through trace-level training data and the 27B model.
44k-record training pool with a strict validation contract. Leave-one-dataset-out splits with leakage rules.
Question isolation verified, calibration pipeline built, external benchmarks run with frozen thresholds.
Training on the selected data pool, adding trace-level supervision, and comparing against specialist baselines.
LoRA adapter, head and model card on Hugging Face, plus a drop-in inference server for agent stacks.
We're looking for teams who will try CyberJev on their own agent traces and tell us where it breaks.