Something happens to you in a game. Not a gray-area thing. Someone says something to you that would end a conversation anywhere else — at work, at a dinner party, in a group chat. You want it on the record. So you open the report menu.
In one game, you struggle to find a reporting category that fits what you want to report. You read the list twice, looking for the thing that actually happened. It isn’t there. You close the menu.
In another it’s there, technically, but first you have to say whether this happened in text chat or voice, and then choose between “offensive language” and “offensive content” and you honestly don’t know, because a person said a thing to you and neither of those phrases describes that.
In a third, you are met with a flurry of options. Too many, in a way. Was that hate speech? Harassment? A threat? It was arguably all three. You pick one and hope somebody downstream sorts it out.
Three games. The same words said to you. Three different answers about what just happened (and in one of them, the answer is that apparently nothing worth reporting happened at all).
The obvious read on my kind-of-but-not-really-hypothetical scenario is that the industry hasn’t standardized their reporting tools. Which is true, but kind of lets everyone off the hook as “no shared taxonomy” sounds like a coordination problem nobody in particular is responsible for. But the longer I’ve spent thinking about this, the more I think this is not that because before you get to whether two companies agree with each other, there’s a stranger question sitting there:
Does even any single company agree with itself?
Open a studio’s code of conduct. Then open that same studio’s report menu. Read them side by side. In every case I’ve looked at, those two documents do not align.
Nearly every studio and platform (that I have observed) has a community code of conduct more detailed than its report menu. That is, their policy routinely names harms the reporting menu can’t express, either because there’s no button at all or because one button is holding several categories.
Three examples:
(do note that these systems get updated constantly, so treat this as a snapshot rather than a permanent record and the pattern is the point, not any one company)
Rockstar’s community guidelines name violent extremism explicitly, under its own heading: glorification or promotion of real-world terrorist, extremist, or criminal organizations and their ideologies. But Red Dead Online’s in-game report menu offers four options: Cheating, Abusive Chat, Disruptive Behavior, Offensive Content. There is no way to report the thing the policy specifically prohibits. Unless it is offensive content? Or maybe it is disruptive behavior? You see the problem here. The commitment exists in the document and nowhere in the pipeline that would let anyone act on it.
2K’s code of conduct prohibits content that’s illegal, offensive, lewd, or otherwise harmful, and then names eight things: nudity, profanity, racism, hate speech, misogyny, self-harm, exploitation, abuse. NBA 2K’s report menu has four buttons: Offensive Language, Offensive Content, Boosting, Cheating/Exploiting. You pick a category, pick a name off a list of everyone in proximity, and the report fires. No confirmation step, no detail field. Also, it is notable that every one of those eight harms listed in their policy is a social harm but their menu only has two buttons that could plausibly hold them.
Blizzard’s in-game code of conduct names hate speech and discriminatory language as inappropriate, and says threatening or harassing another player is never acceptable regardless of the words used. Four distinct things, plainly separated. Overwatch 2’s report flow offers a handful of categories, each with subcategories underneath: Inappropriate Communication, Inappropriate Name, Cheating, and Gameplay Sabotage (other sources also indicate they have an Emergency category but I have not been able to confirm that is new or old information?).
Cheating separates hacking from boosting. Behavior separates going AFK, intentionally feeding, blocking team progress, and griefing. There are four named ways to be a bad teammate. However, Inappropriate Communication separates text chat, voice chat, and spam. Two of those three are not behaviors. They’re channels. Hate speech, a threat, and sustained harassment (three things the policy explicitly distinguishes) arrive as the same object, sorted only by whether they were typed or spoken.
This is a studio that builds precise categories when it decides a harm warrants them. It did it six times over for different ways of playing badly. But the one place it didn’t is the harm players encounter most.
In every case the policy is more sophisticated than the instrument that is meant to implement it.
So far I’ve been describing what gets left out - the policy names more harms than the menu can hold, and things fall off in the translation. If that were the whole story, the fix would be easy… just add more options.
But this is not the whole story. The harms that fall off aren’t a random selection and the ones that survive don’t all survive equally intact.
Jiang and colleagues (2026) ran an automated analysis of codes of conduct across 9,500+ multiplayer titles on Steam. Only about 350 (roughly 4%) had a retrievable, valid code of conduct at all, with availability concentrated among high-profile, well-resourced titles averaging ten times more player reviews than the typical multiplayer game. And among the codes that did exist, gameplay-mechanic violations showed the highest levels of specificity, while interpersonal harms (harassment, hate, discrimination) showed the lowest.
This means there are actually two mismatches running at once:
A vertical mismatch (breadth): This is the loss between layers. That is, the code of conduct names more harms than the menu can express. Eight becomes four. Resolution lost in translation from document to instrument.
A horizontal mismatch (depth). This is the loss inside each layer. That is, gameplay violations get described more precisely than interpersonal ones. That’s true of the policies and also true of the reporting menus.
Put them together and you get the shape of the problem. Interpersonal harm starts with less detail in the policy, and then loses more of it in the compression to the menu. It gets degraded twice. Gameplay harm arrives largely intact, because it had more detail to begin with and because the menu was built to hold it.
Using NBA 2K as an example again. Half of their four-button reporting menu is competitive integrity, and the eight social harms named in the policy collapse into the two buttons left over.
This unevenness isn’t just about being untidy. Romhanyi, Wells and Steinkuehler (2026) surveyed 602 young adult players (across games) and found that almost everyone had seen harm in games — 98.3% reported witnessing toxicity in online games, and 93.9% reported being targeted by it. The most frequent items, for bystanders and victims alike, were trolling and griefing, offensive names, purposeful embarrassment, harassment. So the everyday stuff is not the long tail. It is the distribution (while doxing, stalking, and swatting sat near the floor of the scale). And it’s precisely the material that arrives at the menu already thin in the policy and then gets compressed into “Offensive Language” and “Offensive Content” — two buttons carrying the highest-volume harm category in the game, while competitive integrity gets half the menu to itself.
That’s the volume argument. But it is the norms argument that worries me more.
Social harm isn’t just the most common thing that happens in your game. It’s the thing that teaches everyone what your game is. And the same study found something uncomfortable about how that teaching works: bystander exposure was associated with only one of the four responses the study measured — social withdrawal. Not confronting. Not reporting. Not leaving. Just going quiet. Witnessing toxicity didn’t push people toward doing something about it; it pushed them out of the conversation while the game carried on around them. The authors are careful (this work is cross-sectional, so read it as a pattern rather than a causal chain) but their reading is that repeated witnessing may reinforce the bystander effect rather than erode it. Watching it happen teaches you to stay quiet the next time it happens.
Now put the report menu into that picture. The bystander who goes quiet is a bystander who found no other available move. They witnessed something, and the two channels available to them were confront the person themselves (which the data says most people won’t) or file it. And filing it means opening a menu that either doesn’t name what they saw or offers a bucket so vague it reads as a shrug.
The menu is the low-cost, low-risk action available to a witness. It’s the one form of bystander intervention that doesn’t require standing up in voice chat and taking the heat. When it can’t express what happened, the only remaining option is the one we already know people default to.
Which means a menu that compresses social harm isn’t just under-measuring it. It’s removing the cheapest available alternative to silence, in the exact category where silence does the most damage to the culture of the space.
Community codes of conduct are effective proactive interventions for harms in games (we will talk about that more in a minute). And, louder for those in the back, your report menu is your code of conduct (or at least, it is supposed to be).
A community code of conduct is a policy document; the report menu is the enacted version of that code. Every option in that dropdown is a behavior your company has decided is worth naming, worth routing, worth acting on. Every harm your policy names that isn’t in the dropdown is a sentence that didn’t survive. It exists in the document and nowhere in the machinery.
Codes of conduct get studied as governance - norm-setting devices, enforcement scaffolding, public statements of values. Reporting flows almost never make it into that conversation. They should, because the flow is where the code of conduct actually gets implemented.
Three reasons that reporting flows may be the most consequential governance surface a studio owns.
This year I published a paper that outlines the EFFECT Framework, which are guidelines for developing effective codes of conduct. In this framework I put accessibility (Easy to find and understand) first, because nothing else in a code of conduct matters if nobody reads it. Accessibility fails in two independent ways: cognitively, if people can’t understand it once they find it, and physically, if they can’t find it at all.
Codes of conduct tend to fail the physical test…. pretty badly. Böhme and Köpsell (2010) tracked more than 80,000 users and found people spend about eight seconds on EULAs before clicking through. Grace and colleagues (2022) found that of 60 popular multiplayer games, 22% had no code of conduct on their website and another 32% buried theirs inside EULAs and terms of service. And unlike privacy policies, which reliably sit in a header or footer, codes of conduct have no standard placement at all.
But nobody skims a report menu. You open it at the exact moment something has happened, you’re motivated, and you read every line because you need one of them to fit. It’s the one governance surface players navigate to deliberately, unprompted, at the moment they care most. You cannot buy that kind of attention. We already have it and we’re not using it as effectively as we could be.
The menu is an injunctive norm — a statement of what this community holds should not happen, delivered in one of the few formats players actually read. When it offers five named subtypes of cheating and one vague bucket for everything social, it communicates something quite precise: we take this seriously, and we’re vague about that.
The absence teaches too. A player who opens the menu and finds nothing that fits learns that what happened to them isn’t the kind of thing this place recognizes. That lesson sticks, and it generalizes to the next game and the one after that.
Two findings that should change how studios budget. In work on Minecraft servers, community codes of conduct predicted toxicity rates better than top-down moderation alone, even across different communities on the same platform. And Fang and colleagues (2023) expected more rules to suppress participation and found the opposite: values-centered rules were associated with a 79% increase in user interaction, content-restriction rules with 57%, and comprehensive rules with no drop at all. Restrictions get read as clarifying rather than constraining. People contribute more when they know where the edges are.
Which makes a menu that names harm clearly a behavioral intervention, not just a measurement instrument. Low lift, proactive, reaches everyone.
The other thing we need to keep in mind is that somebody is already absorbing the gap between your community codes of conduct and reporting tools… the digital leadership.
This is the piece that got me here in the first place.
Community managers and moderators are the frontline stewards of everything above. Earlier this year, I published a paper called The CHEER Framework: A New Approach to Assessing Digital Community Health in Games (apparently, this year is the year of the acronym for me). In this paper, I argued that digital leadership maps onto all three pillars of community health and should be treated as co-interpreters of the data rather than just collectors of it. In the DLC Leadership Program, we’ve been building training for this group on the premise that community and safety leaders are the undertrained third pillar of the trust and safety stack.
What that work keeps surfacing is when the menu doesn’t match the policy, digital leaders absorb the contradiction. They’re who a player comes to after finding nothing that fits. They’re the one who has to say “I know, there’s nowhere to put that.” Player expectations research bears this out from the other side as players don’t want leaders telling them what to do, but they very much expect them to keep the space safe. Hard mandate to deliver when the reporting instrument silently disagrees with the values you’re upholding.
Now at this point if you are still asking “But does the menu really matter that much?”... I have one last objection worth taking seriously. Frumkin and Cahill’s analysis of Call of Duty’s player safety work (devcom 2025) found roughly 79% of players flagged by proactive detection had no associated player report at all, and only about a quarter of player reports contained actionable evidence. Player reporting is a small, noisy channel that misses most violations and offloads the documentation burden onto people who were just harmed. If you want the intervention that catches the most bad behavior, it isn’t the dropdown. But detection yield is the wrong measure. Proactive detection finds what it was built to find. Player reports are how you learn about the harms nobody has built a classifier for yet, and the menu is the label space every downstream system inherits. And it’s the only place in the entire stack where the person harmed says what happened, in their own words. Everything else in trust and safety is inference about players. This is the one input that’s testimony from them, which means what the menu can’t express, they can’t say.
If the menu is the enacted code of conduct, then the fix isn’t “add more options.” It’s making the two documents say the same thing. I won’t claim to have all the answers, but I’m happy to give us all a place to start.(My suggestions here drawn from the conceptual work done by many others in the space, including the Thriving in Games Group (formerly the Fair Play Alliance) Disruption and Harms in Online Gaming Framework)
Please note that these categories are not listed in any particular order and it may make more sense to have them listed in one way or another depending on your game.
A proposal like this can read as a demand for uniformity, and that is NOT what I’m arguing. The goal is not uniformity. What makes this idea workable is a fixed core and a flexible layer.
The core encodes harms that hold regardless of genre. A death threat is a death threat in a farming sim and in a tactical shooter. Grooming is grooming in a racing game and in a sandbox.
The extension layer is where genre reality lives. Gameplay sabotage matters in Dead by Daylight in a way it doesn’t in a co-op crafting game. Intentional feeding is meaningful in a MOBA and meaningless in a battle royale. Studios should absolutely build with this in mind as a taxonomy that can’t express what actually goes wrong in your game is one your players will route around.
What makes it work is that extensions nest underneath core categories. They never replace one and never sit beside it at the top level. “Intentional feeding” is a child of Competitive Integrity. “Taking the game hostage” is another child of Competitive Integrity. The core meaning stays stable no matter how much genre-specific detail you pile underneath.
This is the same principle that runs through everything I’ve argued about codes of conduct: they work best when tailored to the mechanics, demographics, and interaction patterns of a specific community rather than applied as a universal template. Context matters. It’s in the framework for a reason.
But context-specific doesn’t mean everything is negotiable. There is a floor, and we already agree on where it is. Nobody in this industry is arguing that grooming is fine, or that credible threats are fine, or that recruitment for violent extremism is fine. We wrote it into our codes of conduct. We say it out loud at GDC every year but haven’t necessarily wired it into the thing people actually use.
There is one other thing that I would add that is not deliberately not on that list, but rather sitting adjacent to it… that is, a category along the lines of “I’m worried about someone.” A separate, visually distinct, non-punitive path for self-harm and crisis concern. It doesn’t go to enforcement. It goes to a wellbeing workflow. Someone telling another player to kill themselves and someone saying they want to die. The first is abuse. The second is a person in crisis. GTA Online files “Threats of Suicide or Self-Harm” as a sub-type of Abusive Communication, which is a framing in which the taxonomy has quietly decided the person in crisis is a rule-breaker. That decision determines which team sees it, how fast, and whether what comes back is a resource or a sanction.
Roblox shows one way to handle it. Their Community Standards give “Suicide, Self Injury, and Harmful Behavior” its own named category, sitting alongside child exploitation and violent extremism rather than nested under abuse. And since the July redesign, filing a report can immediately surface guidance and helpline links while the report is still in the queue. Support and enforcement run on separate clocks, which is right, because the person who needs the helpline isn’t waiting on a moderation verdict.
Two of the categories that I propose above as core categories, Violent Extremism & Terrorism, and Real-World Criminal Activity, may feel out of place to some people. The objection isn’t about design. It’s something like “ Isn’t that a bit much for a game?” or “Do we even need to collect it?” (I’ve heard both said to me).
The short answer is, yes. And also, you are probably already obliged. The DSA requires platforms to give people a way to flag illegal content and act on it. The Online Safety Act sets duties around priority offenses, terrorism among them. Applicability depends on your service and where your players are, but the direction of travel isn’t ambiguous.
There has been plenty of research from groups like the Extremism and Gaming Research Network, the ADL, and various other researchers (including my own work), that supports and justifies why we should have a reporting category for extremism/terrorism in games.
Some platforms address terrorism/extremism explicitly in their policies and have already built in a related type reporting flow into their systems (e.g., “dangerous or illegal activities” category of some kind). So whatever the objection is here, it isn’t practicality.
Also, do notice what the pull-away-this-feels-uncomfortable instinct is telling you. Nobody looks at five named subtypes of cheating and asks whether that’s a bit much for a game. The intuition that aimbot-versus-wallhack granularity is proportionate while a terrorism category is excessive — that intuition is the horizontal mismatch, showing up as a gut feeling rather than a taxonomy. It’s the clearest evidence I can offer that the asymmetry isn’t a resourcing accident. It’s a sense of what we have decided is normal (and/or acceptable) in our gaming spaces. Good news here is we have the power to change that.
I should clarify one thing about how these categories get surfaced. A category labeled “Illegal Content” is not going to be the most effective way forward as it is asking players to assess the law before they can file, and it becomes a dumping ground where harassment, doxing, gambling, drugs, fraud and piracy all land together. But a category that describes observable behavior in plain language (such as someone is recruiting for a violent group) asks nothing of the sort. The system itself can work out what that means legally, in which jurisdiction, and where it goes. A category can be named for a behavior that happens to be illegal but it should not be named for its legal status.
Last but not least (I promise I’ll get off this soapbox in a minute) but extremism earns its own slot rather than sitting under threats, because recruitment isn’t threatening in shape. It’s often warm, welcoming, and patient, which is exactly what makes it work. Filing it under “threats” means the menu is looking for the wrong kind of thing entirely. It’s also the category most likely to be missing.
A shared core isn’t a shared list of words. It’s a shared list of words with definitions attached, because the words alone do far less work than they appear to.
“Harassment” is the obvious case. Two studios can both put it in their menu, both act in good faith, and be counting different behaviors. So the unit that gets standardized isn’t the label. It’s the label plus:
A one-line definition in plain language (what this covers)
The boundary (what it doesn’t cover)
Worked examples on both sides of that line
Xbox has already done this, and it’s among the best I’ve seen. Their Community Standards define harassment against its nearest neighbor (i.e., trash talk is banter about the game that keeps competition healthy; harassment is behavior that’s personalized, disruptive, or likely to make someone feel unwelcome or unsafe). Then they draw the line explicitly (trash talk doesn’t include threats, real-life intimidation, or personal insults based on identity) and show examples of both. Bungie’s expanded Community Standards do something similar throughout.
They’re genuinely good. But again, live in a policy document, on a website, in a section most players will never open. They never make it into the thing people actually read.
Definitions also make the horizontal mismatch I mentioned earlier impossible to ignore. Sit down and write a one-line definition, a boundary, and two examples for every category in your menu. If the cheating definitions come out crisp and the harassment definition comes out as a gesture at a feeling, you’ve found the asymmetry in your game or platform.
In the end, the way we approach our reporting categories and definitions should be living community tools rather than static legal documents. Every taxonomy is a snapshot of the harms we understood when we wrote it, which makes it structurally out of date with respect to whatever is happening right now. The gap between a harm emerging and a category existing for it can be a gap, which could be measured in literal years, because the category gets added after the harm becomes a news story, a regulatory inquiry, or a lawsuit. By then it isn’t emerging. It’s established.
“Other” shortens this window if we treat it as an instrument rather than a sinkhole where things get lost and never come out.
We should be tracking the “other” as a share, not a count. A rising proportion means the taxonomy is drifting away from lived experience, and it rises before anyone can tell you what the new thing is. Also, a spike in Other paired with a dip in an adjacent category often means a new harm is being misfiled into an old label, or a label’s meaning has shifted underneath you.
Also consider clustering the free text on a fixed cadence. Quarterly, minimum. You’re not looking for individual reports; you’re looking for a cluster that didn’t exist last quarter. Also, set a promotion threshold in advance. If a cluster persists two consecutive quarters, it becomes a candidate subcategory. Deciding the rule ahead of time keeps it from becoming a recurring debate about whether the thing is real yet.
Lastly, escalate across titles. The lack of consistency in reporting categories (even core ones) across titles from the same studio is honestly… baffling. The same cluster appearing across multiple titles in a portfolio isn’t a title problem. It’s an ecosystem-level shift, and the highest-value early warning signal in the system (remember ecosystem health is a core component of understanding our spaces).
And when a new harm gets promoted into the menu, it goes into the code of conduct too. That’s the loop. Otherwise you’ve recreated the mismatch in the other direction, a menu catching things your policy never claimed to care about.
I’m not a designer. But I am a researcher, and researchers are good at noticing patterns. Here are some worth considering when aligning your code of conduct with your reporting categories:
Behavior first, surface second. Ask what happened before you ask where. Channel (voice, text, avatar, UGC, DM, gameplay) is metadata the system captures automatically or asks second. Never a gate the player passes through first, and never a substitute for naming the behavior.
Players never make legal determinations. The player describes the behavior. The system handles the law.
Non-criminal harms are also a priority. Harassment, cheating, and griefing aren’t lesser categories to be swept into “other.” They’re the overwhelming majority of what players actually encounter, and they get top-level positioning.
Granularity parity. If you have five named subtypes of cheating and one bucket for everything social, that’s a statement about what you value.
Every category on every surface. If you can be harassed in voice, you can report harassment in voice, with the same label you’d use in text.
The menu must be able to express the policy. Read your code of conduct. Count the harms it names. Open your report menu and count the options. If the second number is smaller, the difference is the set of commitments you made in public and built no way to act on.
Also, I know that earlier I offered ten top-level categories and that may be more than fits comfortably on your UI. If so, you can consider an adaptive flow that skips irrelevant questions. This way, the average path gets shorter while the underlying framework gets richer.
And remember that none of this is fixed in place. Report menus can be changed, they do get changed, and should be revisited and updated.
One example: Roblox spent a long time with an in-game flow that gated you behind a “Type of Abuse” question offering two options (text chat or avatar) while the reasons listed underneath included things like cheating and username, which are neither. A developer laid it out with screenshots on Roblox’s own developer forum in May. By then the fix was already underway. Roblox had announced a reporting redesign in their February Safety Snapshot , which included simpler language, a step-by-step flow, and improved screenshot submission. Their worked example was replacing “profanity” with “swearing.”
On July 1st their Chief Safety Officer published the next stage. Players may not know which Community Standard was broken, they wrote, “and they shouldn’t have to.” The new flow adapts its questions to the type of report and skips the irrelevant ones, and reports now generate a notification when they result in action. A further update in Q3 will let players follow up on reports they’ve already filed. Notably, the flow’s plain language was pulled from Roblox’s Youth Guide to Community Standards, co-designed with their Teen Council.
A large platform made the instrument speak the policy’s language, in stages, with receipts.
So now that you’ve made it through all of these words, how do you get started?
Open your own code of conduct and your own report menu side by side. Count the harms named in one. Count the options offered in the other. Then write a one-line definition and two examples for each.
(also, while you are there maybe audit your community code of conduct too)
You can do that this afternoon. You don’t need anyone’s permission, you don’t need a budget, and you don’t need the rest of the industry to move first. You just need to be willing to look at the answer.
Then make your portfolio consistent. If a single publisher ships titles in the same year where one has a full child-safety taxonomy, one has none, and one has four buttons and no confirmation dialog, that’s an internal alignment problem and it’s solvable this quarter.
The categories I offer above are just a starting point, not scripture. Argue with them. Tell me where the seams are, tell me what your moderation queue sees that this doesn’t hold, tell me if nine is better than ten.
Just remember that your report menu is your code of conduct. It’s the version with consequences attached, the version players actually read, and the version that quietly teaches everyone what this community thinks is worth caring about.
Most studios are running two documents that disagree. The good news is that’s not a coordination problem, or a standards problem, or somebody else’s problem.
It’s a dropdown.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.