❌

Normal view

Towards CSI: What's the best harness? (arXiv 2026)

We studied a question that receives surprisingly little attention:

Does the agent harness matter as much as the underlying LLM?

We benchmarked five different cybersecurity scaffolds while keeping the model fixed (alias2-mini) across all 33 CyBench challenges.

Key findings:

  • No single scaffold performs best across every challenge.
  • Combining heterogeneous scaffolds consistently improves coverage.
  • A shared blackboard architecture solves 19/33 challenges (57.6%), outperforming every individual harness while reducing execution time.

Paper: https://arxiv.org/pdf/2605.28334

Happy to answer technical questions or discuss the benchmarking methodology.

submitted by /u/Obvious-Language4462
[link] [comments]
❌