Evaluation protocol

ReAgent Cyber Benchmark Evaluation Protocol

Status: evaluation protocol. Worthify has not published a verified ReAgent lift result. Any future result must keep model, tool-system, and human-workflow comparisons separate and identify the tested task, budget, environment, and claim gate.

Buyer questions

Questions buyers ask.

Which cyber benchmarks are relevant to ReAgent?

Cybench contributes reverse-engineering CTF tasks. CyberSecEval contributes malware-analysis tasks. CrackMeBench contributes agentic binary tasks with executable validation. SRE-Bench defines a contamination-resistant software reverse-engineering benchmark; task availability must be confirmed before use. Worthify will report each suite and selected task subset separately because the tasks and scoring differ.

Does a higher ReAgent score mean the underlying model improved?

No. The defensible comparison asks whether the same model completes the same held-out tasks more successfully when ReAgent is available. ReAgent changes the tool system around the model; it does not change the model weights.

How would Worthify measure ReAgent for human analysts?

A future human study would compare each analyst with and without ReAgent using counterbalanced, difficulty-matched tasks with no task repeated across conditions. Blinded reviewers would measure time to a validated finding, false leads, corrections, and review work.

Has Worthify published verified benchmark lift?

Worthify has not published a broad, statistically verified ReAgent accuracy-lift claim. Any future result must name the benchmark, models, task count, baseline, budget, effect size, uncertainty, failures, and cost.

Positioning

Keep the claim specific and reviewable.

Keep each benchmark result separate

External suites provide a reference point. Held-out reverse-engineering tasks reduce the chance that a model recalls a public answer. A result names the exact suite version and selected task subset.

  • Cybench reverse-engineering CTF tasks
  • CyberSecEval malware-analysis tasks
  • CrackMeBench and held-out binary tasks; SRE-Bench when its task set is available

Require a controlled model comparison

Before publication, the comparison must freeze the model and tool versions, prompt objective, binary, scoring oracle, time limit, spend ceiling, repeat policy, and exclusion rules. The treatment changes ReAgent access and records actual treatment use.

  • Credible command-line and disassembler baseline
  • Matched ReAgent treatment arm
  • Exact or executable scoring before qualitative review

Require cost and failure accounting

Before publication, the result must retain the full resource envelope and every attempt. More completed tasks do not establish a better system when the treatment consumes far more time, context, or compute.

  • Attempted, completed, correct, incorrect, and excluded counts
  • Wall time, tokens, tool calls, compute, and cost
  • Timeouts, infrastructure failures, and treatment non-use

Define the statistical claim gate

The protocol uses exact-oracle pass rate as the primary LLM endpoint. The pilot sets the powered task count. Paired analysis reports effect size and confidence intervals, clusters related questions by binary, corrects multiple comparisons, and retains every exclusion.

  • Model, harness, prompt, task-manifest, and environment versions
  • Paired binary analysis plus bootstrap uncertainty for continuous metrics
  • Versioned protocol, task manifest, scoring code, run records, and limitations

Design the human study separately

A future human study would use a counterbalanced within-subject design, matched tasks, and no repeated task across conditions. The publication record would include cohort size, experience, training, reviewer blinding, validation rubric, and inter-rater agreement.

  • Time to first independently validated finding
  • False leads, analyst corrections, and reviewer time
  • Tool training and task order reported with the result