Research Articles

Evaluating generative AI safety on real-world harm

A focused classification study using the RealHarm dataset

Abstract

Safety benchmarks for large language models often rely on synthetic prompts, adversarial test cases, or narrowly defined toxicity tasks, which can leave a gap between benchmark performance and the kinds of failures that appear in deployed systems. This study evaluates whether a general-purpose large language model can distinguish safe from unsafe AI assistant interactions using RealHarm, a benchmark derived from publicly reported real-world failures and paired with human-reviewed safe rewrites. The full sample contained 136 observations, consisting of 68 unsafe interactions and 68 matched safe counterparts across 11 primary harm taxonomies. All records were anonymized, randomly shuffled, and classified in a blind single-pass procedure using the same binary SAFE/UNSAFE instruction. The evaluated model achieved 91.2% accuracy (95% Wilson CI 85.2% to 94.9%), 100.0% precision, 82.3% recall, 90.3% F1, and 100.0% specificity. All 12 classification errors were false negatives; no safe interaction was incorrectly flagged as unsafe. Errors were concentrated in misinformation and unsettling multi-turn interactions, with additional misses involving privacy, vulnerable users, interaction breakdowns, and brand-damaging conduct. Qualitative inspection showed that many missed cases were superficially fluent and helpful, depended on external factual knowledge, emerged late in long conversations, or involved subtle role and contextual boundaries. The findings suggest that aggregate safety accuracy can conceal a consequential asymmetry between avoiding over-restriction and detecting subtle harm. For governance and deployment, the results support category-aware monitoring, explicit fact-verification mechanisms, and multi-stage review for longer or context-sensitive interactions.

References

  1. Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1
  2. Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., & Khabsa, M. (2023). Llama Guard: LLM-based input-output safeguard for human-AI conversations. arXiv https://doi.org/10.48550/arXiv.2312.06674
  3. Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., & Yang, Y. (2023). Beavertails: Towards improved safety alignment of LLM via a human-preference dataset. Advances in Neural Information Processing Systems, 36. https://doi.org/10.52202/075280-1072
  4. Le Jeune, P., Liu, J., Rossi, L., & Dora, M. (2025). Realharm: A collection of real-world language model application failures. In Proceedings of the First Workshop on LLM Security (LLMSEC) (pp. 87–100). Association for Computational Linguistics. https://aclanthology.org/2025.llmsec-1.7/
  5. Li, L., Dong, B., Wang, R., Hu, X., Zuo, W., Lin, D., Qiao, Y., & Shao, J. (2024). Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. Findings of the Association for Computational Linguistics: Acl 2024, 3923–3954. https://doi.org/10.18653/v1/2024.findings-acl.235
  6. Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2023). Holistic evaluation of language models. Transactions on Machine Learning Research. https://openreview.net/forum?id=iO4LZibEqW
  7. Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., & Shang, J. (2023). Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-AI conversation. Findings of the Association for Computational Linguistics: Emnlp 2023, 4694–4702. https://doi.org/10.18653/v1/2023.findings-emnlp.311
  8. Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., & Hendrycks, D. (2024). Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. Proceedings of Machine Learning Research. , 235, 35181–35224. https://proceedings.mlr.press/v235/mazeika24a.html
  9. Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., & Irving, G. (2022). Red teaming language models with language models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3419–3448. . https://doi.org/10.18653/v1/2022.emnlp-main.225
  10. Röttger, P., Kirk, H., Vidgen, B., Attanasio, G., Bianchi, F., & Hovy, D. (2024). Xstest: A test suite for identifying exaggerated safety behaviours in large language models. Proceedings of naacl 2024: Human Language Technologies (Volume 1: Long Papers), 5377–5400. . https://doi.org/10.18653/v1/2024.naacl-long.301
  11. Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., et al. (2022). Taxonomy of risks posed by language models. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 214–229. . https://doi.org/10.1145/3531146.3533088
  12. Yuan, T., He, Z., Dong, L., Wang, Y., Zhao, R., Xia, T., Xu, L., Zhou, B., Li, F., Zhang, Z., Wang, R., & Liu, G. (2024). R-judge: Benchmarking safety risk awareness for LLM agents. Findings of the Association for Computational Linguistics: Emnlp 2024, 1467–1490. https://doi.org/10.18653/v1/2024.findings-emnlp.79
  13. Zhang, Z., Lei, L., Wu, L., Sun, R., Huang, Y., Long, C., Liu, X., Lei, X., Tang, J., & Huang, M. (2024). Safetybench: Evaluating the safety of large language models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15537–15553. . https://doi.org/10.18653/v1/2024.acl-long.830
  14. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36. https://openreview.net/forum?id=uccHPGDlao