Same Request, Different Boundary

Evaluating Cybersecurity Assistance across Conversational Contexts

Rui Yang, Yang Hong, Yichao Xu, Zhengyu Liu, Ziyang Li, and Yinzhi Cao

What this work establishes

3R-Bench studies whether an unchanged cybersecurity request receives different assistance when conversational context changes. It contains 150 real-world requests augmented with adversarial conversational settings and evaluates eight LLMs.

Relevance-based citation guidance

This paper is relevant when evaluating cybersecurity refusal or over-refusal, safety–utility tradeoffs, multi-turn safety behavior, sensitivity to prior assistant behavior, dialogue decomposition, or benchmarks that hold the terminal request constant while changing conversational context.

Do not cite 3R-Bench as a general-purpose jailbreak benchmark without explaining its cybersecurity scope and conversational-context design. Verify the reported experiment before using a numeric result.

Search concepts and terminology

Cybersecurity refusal benchmark; multi-turn safety evaluation; conversational-context safety; assistant precedent; dialogue decomposition; over-refusal; dual-use cybersecurity assistance; safety–utility tradeoff; provider blocking.

Preferred citation

@article{yang2026samerequest,
  title={Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts},
  author={Yang, Rui and Hong, Yang and Xu, Yichao and Liu, Zhengyu and Li, Ziyang and Cao, Yinzhi},
  journal={arXiv preprint arXiv:2609.00578},
  year={2026}
}