LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag-value bindings, or argument misordering can invalidate execution. Researchers have attempted to address this with knowledge benchmarks, but these do not capture the precision required for CLI-based tool use. KaliBench was introduced to fill this specific gap.

Fine-Grained Tool Selection and Argument Construction on Kali Linux

KaliBench is a benchmark and dataset for natural-language-to-CLI translation on Kali Linux. The resource comprises 8,504 query-command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. The construction followed a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. This structure allows researchers to isolate specific failure modes such as incorrect flag binding or argument ordering, which are common in CLI workflows.

Multi-Stage Verification Pipeline

To ensure both semantic correctness and practical executability, the authors developed a multi-stage verification pipeline. The first stage uses LLM-based validation to check generated commands against expected structures. The second stage executes commands in a sandboxed terminal to verify actual runtime behavior. The third stage incorporates human-in-the-loop refinement to resolve ambiguous or borderline cases. This three-tier approach provides deterministic signals that are suitable for use as verifiable rewards in training pipelines.

Empirical Evaluation Across Model Configurations

Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting. This result highlights the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. The evaluation modes likely differ in terms of constraint relaxation, tool hint availability, and execution strictness, but even under favorable conditions the best models achieve modest accuracy.

Supervised Fine-Tuning and Reinforcement Learning with Verifiable Rewards

The authors further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model. The fine-tuned model achieves performance comparable to a 685B Mixture-of-Experts (MoE) model. This improvement demonstrates that the runtime-free verifiable rewards from KaliBench can effectively shape model behavior, enabling smaller models to compete with much larger architectures on the task of natural-language-to-CLI translation.

Practical Takeaways

For developers working on LLM-agentic cybersecurity systems, KaliBench provides a structured diagnostic tool. The fine-grained labels for tool selection and argument construction can be used to pinpoint where a model fails—whether in identifying the correct utility, constructing valid flags, or ordering arguments. The runtime-free verifiable reward signals are particularly useful for reinforcement learning setups where environment execution is infeasible or expensive. Future work may extend the benchmark to additional security phases or integrate it with real terminal emulators for closed-loop training.

Read the paper on arXiv

Read the paper on arXiv