Towards CSI: What's the best harness? (arXiv 2026)
We studied a question that receives surprisingly little attention:
Does the agent harness matter as much as the underlying LLM?
We benchmarked five different cybersecurity scaffolds while keeping the model fixed (alias2-mini) across all 33 CyBench challenges.
Key findings:
- No single scaffold performs best across every challenge.
- Combining heterogeneous scaffolds consistently improves coverage.
- A shared blackboard architecture solves 19/33 challenges (57.6%), outperforming every individual harness while reducing execution time.
Paper: https://arxiv.org/pdf/2605.28334
Happy to answer technical questions or discuss the benchmarking methodology.
[link] [comments]