Tattle conducts periodic data annotation exercises to grow the Uli dataset - a crowdsourced dataset of harmful terms and phrases in multiple Indian languages. As language evolves, so do the harms associated with it, and hence we aim to keep our crowdsourced safety dataset up to date with frequent annotation workshops. Our most recent annotation workshops surfaced interesting findings about how interactions online become unsafe even without the use of slurs, and the pathways for community moderation that have evolved to address the complex arena of online content moderation.

Tattle gathered 19 annotators over multiple workshops with expertise in research, journalism, policy development, monitoring and evaluation, education, social work, and digital activism hailing from organizations in academia, media, and the development sector. There was a mix of people who had supported Uli slur annotation exercises in the past and people who were completely new to the process. We shared 200+ terms/phrases in a range of languages spoken by the annotators including Hindi, Malayalam, Tamil, Bengali, Assamese, Konkani, Marathi, and Gujarati. The last four were new additions to the dataset bringing the total languages represented in the Uli database to 11. We asked annotators to fill in metadata for predefined categories for the given slurs, as well as, add new slurs with metadata. This dataset was culled from Tattle’s crowdsourced Uli slur list developed over previous rounds of crowdsourced annotations. Since we did not have an existing database of words for Assamese, Marathi, Konkani, and Gujarati, these lists were created from scratch by the respective annotators.

There were 9 annotators for Hindi/Hinglish, 3 for Malayalam, 4 for Marathi, and 2 for Bengali, and 1 for Assamese. One of the annotators was able to supply slurs in Gujarati, Konkani, and Marathi. Each workshop featured structured time for independent labeling, as well as collaborative discussion regarding the specific annotation categories. These annotation categories were developed for the Uli dataset and included severity of slur, identities targeted, definitions, examples of use, and whether a word had been reclaimed by the community it targeted. Participants engaged in focus group discussions exploring creative solutions for content moderation within India’s digital dating and social landscapes. Finally, annotators also provided Tattle with critical feedback on the design of the exercise itself. In total, the expert group created an annotated database of 400+ harmful terms and phrases in 9 languages.

We posed different questions during the focus groups discussions to understand people’s response to the slur lists as well as their unique experiences encountering online harm. One of our questions garnered extended discussions about the implicit and insidious ways that online harm proliferates - “What makes an online interaction harmful?” Participants reflected on the moment they knew online conversations had turned uncomfortable. One annotator mentioned that often right after they disclosed being a religious minority in direct messaging, they noticed conversations would fizzle or start to feel really awkward. Nothing offensive had been said but the implicit presence of islamophobia was viscerally felt by them. This prompted a gender diverse person to share that similar disclosures are expected of them across online interactions with strangers. If they do not disclose themselves as gender diverse/trans, people accuse them of impersonation or catfishing. But disclosures of identity eventually lead to a similar fizzling of conversation. Both annotators felt this kind of silent awkwardness could be just as disappointing as direct attacks on their identities with slurs. A broader discussion ensued of visual and textual cues that the group felt could invite harassment, awkwardness, and implicit threats in online interactions.

For queer folks and cis women, such implicit threats could occur even before any online interaction had begun. A few people in the workshops discussed that images on their profile pages were subject to excessive scrutiny and often became the first site of harassment. Judgements on how much skin is being shown, assumptions about someone’s sexual availability based on their makeup and clothes, and persistent unwanted sexual advances based on political views were some common trends observed by the group. Participants mentioned that identifying as polyamorous was equated with moral laxness and sexual promiscuity, undermining consent as the basis of any respectful relationship.

From these discussions there was an emerging sense that online harms were often not encoded in just the language used to interact with someone. There was consensus among the trans folks participating in the discussion that being visibly gender nonconforming (such as choosing to showcase oneself in images wearing or not wearing certain clothes for gender euphoria) often made them the target for stalking and sexually violent overtures. These inferences mirror other research findings on the experiences of gender minorities and marginalised communities in online social spaces12.

All contributors agreed that content moderation policies often failed to protect people with intersectional and multiple marginalized identities as the safety of minorities required nuanced moderation ethics and training that would take seriously the widespread prevalence of transphobia/casteism/sexism in society. However, this task is particularly tricky in a context where those developing moderation policies and practices may themselves be implicitly biased towards these groups. Further, if content moderation was strictly enforced to prevent sexual violence and harassment, it might end up removing a large majority of users from online platforms. Therefore, many of the annotators felt the goals of making minorities safer in digital spaces were directly in conflict with platform incentives to keep users on their apps for as long as possible.

These observations reiterate the mundane and hard to pin down nature of online harm. While slurs and hate speech are a predominant issue for anyone navigating online spaces, those with marginalised identities are often subject to subtler threats that are difficult to report or present as evidence of harassment. In an earlier set of annotation workshops, experts had discussed the specific example of a reporter receiving a “Good morning” message from an unknown number every morning to emphasize that they were being watched. Such implicitly coded interactions can often go further in silencing people online than more visible forms of hate owing to the difficulty of enforcing action against them.

While this was a difficult reality of online socializing, instead of moving the discussion into a cul-de-sac it opened up creative ideas around what might be other ways to address online safety that did not rely on centralized moderation or enforcement of ethical practices by platforms. One big way that marginalized folks were already protecting themselves online was by gathering their community together to report content and users collectively. A ciswoman stated that she had once requested friends to report persistent harassers on dating apps to act as a safety measure for all women on the app. The hope was that a series of complaints would reduce the likelihood of this person being visible to other folks in their target demographic. This developed into an intricate discussion on the need for trained bystanders to provide safety online drawing lessons from movement spaces that have created community bystanders for intervening in the place of punitive authorities.

The next topic of discussion was related more directly to the slurs as annotators commented on the types of words they were observing in the database. Since many annotators had been involved with earlier rounds of annotations, they were able to observe some shifts in the lexicon of online abuse. Slurs originating from Hindi had in recent years made their way to Bengali. One expert commented on how certain words though gendered in their origin had become associated with performing specific actions (‘chatukar’; ‘chutiyapa’). This prompted another annotator to observe that in Marathi there were a large number of slurs that specifically targeted the activity of idleness, or that a number of sexualised and gendered terms were used to insult someone as being ‘good for nothing’. These observations are important moments pointing to how language reveals deeper biases that might be shifting over time. Maybe the observation about Marathi slurs could mean that idleness or being unemployed is seen as particularly problematic. On closer examination of the specific slurs themselves, it might reveal that whether someone is implying idleness or unemployment as undesirable, the terms used are all sexualised, indicating that insults are deeply gendered irrespective of their intended use. We are eager to read more from linguistics and historians on how abuse/slurs evolve with language and context.

These rich discussions demonstrate the complex dynamics of reporting abuse in Indian digital space and the subtle ways harm, exclusion and discrimination proliferate beyond the direct use of slurs. As we continue to build the Uli database, we continue to grapple with the many faces of online harm, and innovative responses that we may support to increase people’s agency in the digital commons.

Footnotes

  1. Prarabdh Shukla et al., ‘Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch’, in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ed. Wanxiang Che et al. (Association for Computational Linguistics, 2025), https://doi.org/10.18653/v1/2025.acl-long.1110; David Hartmann et al., ‘Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations’, Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (New York, NY, USA), CHI ’25, 25 April 2025, 1–26, https://doi.org/10.1145/3706598.3713998

  2. ‘Side-Swipe: The Challenges of Online Dating While Trans’, Michelle Sheppard, 29 May 2018, https://mishsheppard.com/2018/05/29/side-swipe-the-challenges-of-online-dating-while-trans/; Molly Grace Smith, ‘Queer Enough to Swipe Right? Dating App Experiences of Sexual Minority Women: A Cross-Disciplinary Review’, Computers in Human Behavior Reports 8 (December 2022): 100238, https://doi.org/10.1016/j.chbr.2022.100238.

Text and illustrations on the website is licensed under Creative Commons 4.0 License. The code is licensed under GPL. For data, please look at respective licenses.