Emnlp26findings
HarmProfile accepted to Findings of EMNLP’26. It characterizes harmful-content distributions in frontier LLMs (23 models, 80K+ artifacts), beyond binary safety scores.
HarmProfile accepted to Findings of EMNLP’26. It characterizes harmful-content distributions in frontier LLMs (23 models, 80K+ artifacts), beyond binary safety scores.