The Manual Problem: Why "How-To" Content Is So Hard to Moderate
Most content moderation is about the message a piece of speech carries: discerning whether a post, an image, a video expresses something harmful. But some of the hardest cases in trust & safety carry a different kind of danger. They're instructions: a step-by-step guide to building a weapon, a detailed description of a self-harm method, a how-to for planning an attack. That content isn't trying to convince anyone of anything or sway popular opinion. It's trying to hand someone a highly transferable skill, and that's what makes it so hard to catch.
Why "How-to" Content Breaks the Usual Rules
Normally, moderation leans on intent and audience: is this post hateful, is it harassment, is it misinformation? Instructional content doesn't play by those rules. The exact same information can show up in a chemistry textbook, a Wikipedia article, a true-crime documentary, a journalist's investigation, and a plan to hurt someone, and it can look almost identical in all five places. A filter that just looks for "the words" ends up flagging legitimate science and news coverage, while someone determined to dodge it can reword things, split the content across several posts, or bury it in slang.
This is often called the dual-use problem: the same information is genuinely useful and genuinely dangerous, depending entirely on who's reading it and why. That means a huge share of the actual decision comes down to context – who's posting, the content itself, where it's shared, concurrent warning signs, overlapping real-world events – rather than only the words themselves.
Three Flavors of the Manual Problem
Weapons and Attack Instructions
Over the past decade, tech companies have built shared infrastructure specifically to catch known dangerous material. The Global Internet Forum to Counter Terrorism (GIFCT) (started by Facebook, Microsoft, Twitter and YouTube in 2017) runs a shared library of "fingerprints" of terrorist and extremist material. The idea is simple: once one company identifies a bomb-making manual, an attack manifesto, or a piece of propaganda, it gets converted into a unique digital fingerprint, not the actual content, and added to a shared database that other companies can check against. If the same content shows up on their platform, they can act on it immediately, without ever sharing user data. By late 2023 that library held roughly 400,000 pieces of flagged content, expanded in 2021 to specifically cover attacker manifestos and the links that point people to them.
The obvious limitation: this only stops things that have already been seen once before. It can't catch something brand new. That's the gap the Christchurch Call was built to close after the 2019 New Zealand mosque shootings: a system where governments and platforms coordinate in real time when an attack has an online component, so everyone can act fast together instead of each company scrambling alone. It's been utilized more than 140 times since 2019, and it worked effectively after the 2022 Buffalo shooting, where the shooter's livestream was pulled quickly and the video and manifesto were fingerprinted and shared industry-wide within hours. But the same organizers admit that the system works much better for attacks that resemble Christchurch (a lone attacker livestreaming from a Western country) and much worse for messier, less familiar situations, like bystander footage from a warzone or large-scale attacks such as bombings. Building a system around one type of tragedy won't automatically generalize to the next one.
Smaller platforms are the weak link. Big companies like Meta, Google, and Microsoft have entire teams and shared tools for this. A small forum or niche app usually doesn't, and bad actors know it, which is exactly why extremist material tends to migrate to smaller, less-resourced corners of the internet once it's been kicked off the big platforms. A nonprofit called Tech Against Terrorism exists mostly to fill that gap: it scans the web for known terrorist content and alerts small platforms directly, and when it does, the vast majority of that content gets taken down within a couple of weeks, providing proof that small platforms can act quickly if there’s someone to point at the problem first.
On the legal side, the law mostly focuses on intent, not information. In the US, it's not illegal to explain how explosives work. That information can be found in libraries, textbooks, and classrooms. What's illegal is teaching or sharing that information when you intend it to be used to hurt someone, or you know the person you're telling intends to use it that way. That distinction is hard to enforce automatically. A piece of software can spot a chemical formula, but it can't tell you what's going on in someone's head.
Self-Harm and Suicide Content
This category runs on completely different logic, borrowed from public health rather than counter-terrorism. Decades of research, summarized by the suicide-prevention charity Samaritans, consistently shows the same pattern: graphic, detailed coverage of how someone died by suicide can lead to copycat behavior in people who are vulnerable, while stories about getting through a crisis and finding help tend to encourage people to reach out instead. Erasing suicide and self-harm from the internet entirely is neither possible nor useful. The actual goal is narrower: strip out the specific, actionable method detail while still leaving room for people to talk openly about their struggles and find support.
Samaritans' guidance for platforms lays out the trade-offs clearly. Automated detection is fast and can catch huge volumes of content, but it's also easy to fool: someone can swap a letter, add an emoji, or use community slang, and a keyword filter will miss it entirely. It also produces plenty of false alarms, which frustrates ordinary users and erodes trust in the platform. Human moderators are much better at picking up on tone, sarcasm, and context. Samaritans gives the example of a photo of a railway line, which looks completely ordinary until you read the caption. Humans can't read every post on every platform, though, and doing this work day after day takes a toll on the moderators themselves. In practice, every serious platform ends up using both: automation to sort the haystack, humans to find the needles.
What gets added to the process is just as important as what gets removed. Samaritans consistently pushes platforms to pair any removal with a signpost to real support – a crisis line, a helpline, something available any time of day – because for someone in crisis, being pointed toward help in that exact moment can matter enormously. This mirrors long-standing guidance for journalists covering suicide: don't describe the method, don't put it in the headline, don't sensationalize it, and always include where people can go for help. Platforms are now essentially following the same playbook newsrooms have used for decades, but at a scale and speed no newsroom ever had to maintain.
AI-Generated Instructions
Generative AI adds a new wrinkle to the problem. In the past, a bad actor had to find a dangerous document somewhere online. Now, in theory, they could just ask an AI model to write one from scratch, phrased however they like. That makes it useless to just search for known bad documents, because the document doesn't exist until someone asks for it.
AI companies have responded by trying to build the safeguard into the model itself, before content ever gets a chance to be posted anywhere. Anthropic, for example, describes a layered approach that includes tailoring what a model will do based on who's using it and in what context, plus real-time filters that check both what a user asks and what the model is about to say back, before it's shown to anyone. The company's Frontier Safety Roadmap also describes ongoing internal work to try to break its own safety filters around chemical and biological weapons, essentially hiring people to trick the system before someone with bad intentions does. Platform moderation catches the manual after it's posted. This approach tries to stop the model from ever writing an equivalent one in the first place. Researchers keep finding new ways to sneak past these filters, though, so it's still a moving target.
Core Commonalities
Looked at broadly, these three disparate scenarios share a few things in common.
Bad actors adapt constantly; the moment a filter or fingerprint system exists, there's an incentive to dodge it: reword it, split it up, misspell it, wrap it in a joke or irony. Every framework above has needed to keep adding new fingerprints and rules over time, because the old ones stop working almost as soon as they're deployed. Laws move slower than any of this: it took the US Congress the better part of a decade to pass a narrowly written law on sharing bomb-making instructions, so platforms are almost always building their own rules in the gap before the law catches up. They can’t rely on speedy government intervention as a solution.
Small platforms carry the most risk with the least resources. Big platforms have teams, tools, and shared databases. Smaller ones usually have neither, and that's exactly where dangerous content tends to end up once it's kicked off the bigger sites. The GIFCT database, one of the most mature tools in this space, isn't open for outside researchers to inspect, because publishing exactly what triggers detection would also hand bad actors a map of how to avoid it. The field has yet to find an easy way to balance this trade-off between transparency and effectiveness.
Automated systems miss context that human moderators understand. A trained human can usually tell a chemistry teacher from someone planning to hurt people, or dark humor from a genuine threat. A machine scanning millions of posts a day mostly can't. Every viable framework in this space includes routing the hardest calls back to a human, rather than pretending a machine can fully replace human understanding.
The Bottom Line
At TrustLab, we’ve yet to see a one-size-fits-all framework that solves this problem outright. What we’ve watched good teams do successfully is slow things down: make dangerous content harder to find, harder to spread, and easier to spot once it's out there, while giving people an escalation point when the content is about their own safety rather than someone else's. The content by itself is rarely the whole story. What determines the right response is context, pattern, and how fast one platform can share what it learns with everyone else.
Sources
- GIFCT — Hash-Sharing Database and Incident Response Framework
- Tech Against Terrorism and its Terrorist Content Analytics Platform
- Christchurch Call Foundation — Crisis Response Protocol
- Samaritans — Industry Guidelines for Managing Self-Harm and Suicide Content Online and Media Guidelines for Reporting Suicide
- Congressional Research Service — Bomb-Making Online: An Abridged Sketch of Federal Criminal Law
- Anthropic — Responsible Scaling Policy and Frontier Safety Roadmap

