Description
This video presents new research on how large language models (LLMs) can unintentionally expose users to potentially radicalizing content. By testing two widely used public models, OpenAI GPT‑4o and Google Gemini, the study shows that even benign prompts about the Tradwives community can bypass safeguards, especially when roleplay or retrieval‑augmented generation (RAG) is involved. The findings suggest that, despite guardrails, LLMs may still generate harmful associations between everyday online communities and extremist narratives, including those linked to violent extremist organizations responsible for severe harm and human rights violations. The video highlights why understanding these risks is essential for researchers, policymakers, and tech companies working on online safety.
Related topic
New Technologies and the Online Dimension
Link
Language:
English
Date:
27/06/2026
Proposed by:
ELIAMEP
Source:
Cyber Threats Research Centre (CYTREC), 2026