Project BroBots is a multi-agent research initiative to identify, classify, and counter toxic online content using NLP-based harm detection and counter-narrative generation. It employs fine-tuned large language models on social media corpora, the Agentic Alliance tech stack, and is motivated by the harms of automated bot-driven misinformation and harassment across internet platforms.
Semantic Classification
Content
Project: BroBots
- To address the negative impacts of toxic online behavior by developing a multi-agent system that can identify and counter harmful content.
- The internet is increasingly populated by bots and trolls that spread misinformation and engage in harassment. This has a negative impact on online discourse and can lead to real-world harm.
- We propose to build a multi-agent system that can:
- Identify harmful content: Use natural language processing (NLP) to identify toxic language, hate speech, and misinformation.
- Counter harmful content: Generate counter-narratives and engage with users in a positive and constructive way.
- Promote healthy online communities: Encourage positive online behavior and create a more welcoming and inclusive online environment.
- Large Language Models: Llama 3 70B, Mixtral 8B
- Fine-tuning: Fine-tune a smaller model on a corpus of Reddit data to identify and classify harmful content.
- Agent Framework: Use the Agentic Alliance tech stack to build and deploy the multi-agent system.
- Data Sources: Reddit API (if available), other social media platforms.
- Data availability: Access to social media data is becoming increasingly restricted.
- Defining “harmful content”: What constitutes harmful content is subjective and can vary depending on the context.
- Ethical considerations: It is important to ensure that the multi-agent system is used in a responsible and ethical way.
- Digital Society Harms
- Death of the Internet
- Agents
- Decentralised Agent Coordination Initiative