Skip to main content
Identrics
Blog Post

How to Automate Hate Speech Detection in Online Communities?

Written by

Published

Updated

Reviewed by Yolina Petrova, PhD

Reviewed according to our Editorial Standards

Online communities face a constant battle against hate speech, fake comments and toxic discussions. Manual moderation is overwhelming and expensive for publishers and forum managers handling thousands of daily comments, which is why more organisations automate hate speech detection.

But hateful comments are rarely isolated. The same messages are repeated, adapted and amplified across communities, often to push a narrative. This guide covers both layers: automating comment-level detection, and seeing the harmful narratives and actor networks behind it with Annex.

Current problem with hate speech in the media

In the final quarter of 2017 alone, Facebook removed 1.6 million pieces of content containing online harassment and hate speech. In the second quarter of 2021, that number jumped to over 31 million hate speech comments. And that’s just one website during a single 3-month period.

Needless to say, online hate speech and disinformation are prevalent, and they are not going away any time soon. This online phenomenon identified in comments exists in all communities, from social media networks and news sites to forums and blogs.

Wherever people can openly speak their minds and practice free speech, negative, fake and hateful content and online hating behaviour can be found.

It is a growing global problem

In many countries, hate speech online is a growing hate crime problem and one that creates unique challenges for media organisations. As we all know, the publishers’ industry is responsible not only for its original content, but also for the content its users generate on their relevant media accounts and websites.

Imagine, for instance, that you run a news website with a comments section that updates in real-time.

You post a news story about migrants or minorities and it goes viral. Great, you now have a story that could be seen by millions. But you will also have thousands of comments to deal with. Those comments could flood your page at a rate of 100s per hour.

But why can this be a concern? Comments are an essential ingredient to a published piece. They can easily shift the original message of the text and make the reader misinterpret the primary meaning. Large amounts of negative comments can change the narrative using expletives and hate speech. And let’s not forget that, you, as a publisher, are responsible for each comment and its content that could cause online harm.

Most often, comments are used to predefine the narrative. Human moderators hardly can read and verify every single comment real-time. This is a vulnerability that is exploited by interested parties.

Comparing online hating based solely on personal beliefs is not feasible, as it is not feasible to check comments manually. And even if you could do that, how would you prevent hateful comments from being posted in the first place? If you’re doing it manually, it means you’ll need to hire numerous people to work around the clock manually checking the comments.

It’s expensive, it’s time-consuming, and it’s impractical.

The result is that your viral news story becomes riddled with hateful comments. Not only will these offend your readers, but online haters could also induce negative attitudes toward and harm the reputation of your brand.

After all, who wants to devote themselves to a community constantly flooded with hateful comments or online trolling? Who wants to spend time on a website that fills with negativity every time a new story is published?

We are not just talking about comments that can be perceived to be mildly inappropriate. They certainly exist, but “hate speech” typically refers to vile and obscene messages filled with online hate speech and vitriol. No one wants to witness online hating when they’re trying to catch up with the day’s news.

Why comment-level moderation is not enough

Comment sections are essential for community and engagement, so removing them is not the answer. The task is to distinguish hate speech and deal with it. Most existing tools fall short in two ways.

First, keyword filters catch obscenities but miss hate speech. A commenter does not need bad language to say something hurtful, and bad language also appears in harmless comments. Human moderators help, but they cannot read everything in real time or know every new slang term.

Second, every tool that judges one comment at a time misses the pattern. When the same framing appears across dozens of threads and outlets within hours, often from a small cluster of accounts, it is no longer individual opinion. It is a narrative being pushed, and it only becomes visible when related content is grouped and connected to the actors behind it.

Effective protection therefore needs two layers: automated detection with human review for individual comments, and narrative-level analysis to reveal coordinated campaigns.

How to automate hate speech detection

1. Choose the right AI moderation tool

Research solutions that support your language and platform (e.g., Identrics’ solutions, Perspective API, custom AI models).

Consider tools with proven accuracy and support for your community’s size.

2. Integrate the tool with your platform

For WordPress, Disqus, or similar, check for ready-made plugins or APIs.

For custom sites, use available REST APIs or SDKs for integration.

Ensure you comply with GDPR and local regulations.

3. Set up moderation rules

Define the thresholds for automatic flagging or hiding of comments.

Set up human-in-the-loop review for borderline cases or flagged content.

Customise filters for specific keywords, phrases, and context.

4. Review and retrain

Routinely review flagged comments for accuracy.

Use moderator feedback to retrain and improve the model.

Update rules and lists based on new trends and emerging hate speech tactics.

5. Monitor, report, and stay compliant

Track moderation outcomes (false positives, undetected cases).

Generate regular reports for transparency.

Ensure your process aligns with legal requirements for hate speech in your region.

Identrics’ hate speech detection solution

At Identrics, we use a human-in-the-loop hate speech detection model.

Tool/MethodAutomation levelLanguages supportedHuman-in-the-loopPricing
IdentricsHighMultilingualYesCustom
Perspective API (Google)ModerateMultilingualOptionalFree
Customer Keyword FilterLowAnyNoFree

Our software checks and flags the comments that may contain hate speech as they are posted. These comments are then sent for human moderation.

The human moderators are directed to the exact words that may contain hate speech, thus allowing them to make sound decisions.

It means that communities can benefit from the ease, simplicity, and speed of automation while still utilising the expertise that only human interaction can bring.

And that’s not all.

If the human moderator determines that the flagged comment is perfectly harmless, they can send it back to the model. The model is constantly learning and improving and grows from a moderator’s feedback on why this comment should not have been flagged, knowing not to flag such a comment in the future.

The longer the model remains active, the more comments it reviews and the more it learns. As it grows, it becomes more effective over time at making these decisions and ensures that fewer false readings are sent for moderation.

Beyond single comments: detecting harmful narratives with Annex

Comment-level detection tells you that a message is hateful. It does not tell you whether that message is part of something bigger. Annex adds the narrative layer.

1. Cluster related content into narratives

Annex groups scattered posts, comments and articles that share framing and claims into topics and narratives, so moderators and analysts see the story being told rather than thousands of separate messages.

2. Track how narratives spread

Narratives are followed across communities and sources over time, showing where a hateful framing started and where it is gaining ground.

3. Map the actors behind them

Annex maps the networks of accounts and outlets amplifying each narrative, helping trust and safety and research teams spot coordinated behaviour that single-comment moderation misses.

4. Annotate and report

Analysts annotate and refine narratives and turn findings into reports, so decisions about moderation, response and escalation are based on the full picture.

Benefits of combining detection and narrative intelligence

Chapter Three of the Bulgarian Criminal Code, “Crimes against the Rights of Citizens”, addresses hate speech and the need to eradicate it. Lawmakers rarely concern themselves with how content gets onto a platform; they expect it to be dealt with.

Automated detection keeps communities safe and supports compliance by catching hateful comments quickly. Narrative intelligence adds context: it shows whether a wave of hate is organic or coordinated, which themes are driving it and who is amplifying it. Together they let organisations respond to the cause, not only the symptoms.

Explore our research on how false information spreads online and AI-generated trolling for more on how harmful narratives are built.

Frequently asked questions

How can I automate hate speech detection in community forums?

Use AI-powered moderation tools that integrate with your platform. Set up custom filters, use human-in-the-loop workflows, and routinely retrain your system for best results.

What are the best AI tools for hate speech moderation?

Popular options include Identrics, Google Perspective API, and custom models built with open-source frameworks. Choose based on language support, accuracy, and integration needs.

Can hate speech detection models work in multiple languages?

Many modern tools support multiple languages. Always confirm that your chosen tool covers your main community language(s).

How do I balance automation with human moderation?

The best approach uses automation for first-level screening, with human moderators reviewing edge cases or appeals. This ensures accuracy and fairness.

Is automated hate speech detection accurate?

Accuracy depends on the model, training data, and ongoing supervision. Combining AI with human review achieves the most reliable results.

How do I know if hateful comments are coordinated?

Look beyond single comments. Narrative intelligence tools such as Annex cluster related content into narratives and map the accounts amplifying them, revealing repeated framing and coordinated behaviour.

Sources

  1. Schmidt and Wiegand, "A Survey on Hate Speech Detection using Natural Language Processing", SocialNLP, 2017
  2. Regulation (EU) 2022/2065 - the Digital Services Act, EUR-Lex

About the author

NV

Nesin Veli

Chief Executive Officer, Identrics

Nesin leads Identrics' work on automation and data transformation. He designs and implements technological solutions that align with the developing needs of the media intelligence market and translates editorial workflows into engineering systems.

Connect on LinkedIn

Reviewed by

YP

Yolina Petrova, PhD

Chief Operations & AI Officer, Identrics