Artificial IntelligenceScience & Tech

LLMs: Models for Language Marginalization

Introduction

Artificial intelligence (AI) can appear objective or neutral on the surface, but its outputs may not be as impartial as they seem. Companies such as Google and OpenAI use Large Language Models (LLMs) in their generative AI services, which can embody the viewpoints of dominant groups in society and amplify ongoing issues surrounding cultural appropriation and ownership. These companies have data sets biased towards Eurocentric views, especially those that are common in the United States (US). Since generative AI creates new content based on preexisting data sets, it can replicate the harmful prejudices that exist in human society. AI can often affirm language biases and propagate racism into the public sphere, exacerbating the social inequalities that minority groups experience in the US.

Text-generating AI is built upon LLMs, which analyze language data to predict the most “human” sounding combination of words to form sentences. LLMs synthesize online resources such as articles, books and forums to learn patterns and echo prominent styles of writing. Certain parameters are put into place by humans so that AI prioritizes certain information over others, which helps it develop a specific voice and perspective. In this way, the selection and extent of online resources inform AI’s accuracy and the information that is propagated by generative technologies. 

However, if the data and those who control the data do not have diverse perspectives, generative technologies may produce inauthentic representations of marginalized groups. ChatGPT, owned by OpenAI, is a notable example. The data given to the LLMs is often limited to prevailing Western perspectives and conventional English speaking sources, and LLMs pick up stereotypes and overarching patterns reflected in these data sets when predicting the probability of words or ideas. This can cause minority dialects or vernaculars such as African American Vernacular English (AAVE) to be misrepresented, and monolithic stereotypes of such minority groups to be perpetuated by what is quickly becoming a pervasive and essential tool.  

LLMs and Dialect Prejudice

As globalization progressed after British colonization, English was adopted as a lingua franca – a universal language – due to Western prominence in business and trade. However, within English, there are linguistic hierarchies, and many minority dialects and vernaculars are devalued compared to “middle-class, white” English. Across the US, those who speak in AAVE are at a disadvantage when advocating for basic rights such as housing. Being and “sounding Black” is unfavourable, as shown in the systemic racism within the US that marginalizes and disenfranchises African Americans. 

This social dynamic is reflected in how generative AI perceives and recreates AAVE. A study by scholars at the University of California-Berkeley reported that GPT-4, one of the recent iterations of ChatGPT, had “greater difficulty” understanding AAVE speakers compared to non-AAVE speakers. Furthermore, when generating texts that aim to emulate AAVE, ChatGPT produced texts that appropriated simplistic and degrading stereotypes about AAVE. This is referred to as “dialect prejudice,” which reinforces the notion that vernaculars such as AAVE are considered unfavourable and incomprehensible. Dialect prejudice also decontextualizes the richness and history of minority vernaculars, which may affect how well speakers understand their identities in later generations. These biases are ingrained within the LLMs, which is problematic as they encourage users to consider “standard” English as the only acceptable version of English.

The lack of digital content created by minority groups in their dialects is a driver in retaining linguistic hierarchies. Minority voices are simply not reflected in data sets, which causes tools like ChatGPT to generate false presumptions based on the assumptions that are expressed by non-minority groups online. For instance, developed countries are overrepresented in the netizen population, and even within developed countries, underrepresented groups (e.g. transgender people and immigrants) are often harassed in online forums for sharing their thoughts. This causes users of underrepresented backgrounds to stop posting content, which lessens the amount of diverse voices in data sets fed to LLMs. 

Even for underrepresented users who post, LLMs have guidelines that may block their visibility in the data sets. Most LLM programmers will include filters to block discriminatory slurs and negative language from being processed by the AI, which is crucial. However, some minority groups have reclaimed traditional slurs into their vocabularies, such as the word “twink” for many LGBTQ2IA+ communities. This gets filtered out by LLMs, which is discriminatory towards their vernacular. Although the intentions may be positive, filters can result in data sets with an extreme overrepresentation of majority groups in developed countries, which is replicated by generative AI.

Positive Effects of LLMs

At the same time, LLMs and generative AI can also positively affect minority communities. LLMs have increasingly been used to maintain endangered languages. For example, speakers of Nheengatu, an Indigenous Amazonian language, have collaborated with LLM researchers to document their language. The LLM researchers made a mechanism that can translate between Nheengatu, English and Portuguese, which helps youth learn Nheengatu without having to know a fluent speaker. This prototype’s data set was made from 7,000 example sentences in Nheengatu. The platform could also generate new sentences using this data set, which helped students learn via conversational methods.

While preserving endangered languages is valuable, the inheritance of languages by technology companies should be treated with vigilance. Widely spoken languages such as English do not have a central “owner,” but endangered languages typically only have several speakers who “own” their language. If technology companies inherit ownership of endangered languages, there may be a risk that lapses in LLM judgements are not caught by them, which creates inauthentic representations of the language. If there is no original speaker to verify the correctness of translations or generated sentences in the endangered language, the “preserved” version may deviate from and be assumed to be the original language.  

Solutions

While LLMs can have positive effects on marginalized communities, there are still measures that should be put into place to ensure the safe and equitable usage of generative AI. Especially for the negative effects regarding discrimination against vernaculars and the appropriation of dialects, concrete solutions are needed to stop these practices. Some proposed solutions are:

  1. Equal internet access. As mentioned above, there is a disproportionate online representation of non-marginalized people in developed countries. Thus, facilitating opportunities for those in developing countries to be represented is crucial. This may take the form of research being done on these nations or creating infrastructure to help citizens obtain internet access to contribute to online discussions. Many underrepresented users in developed countries also experience barriers when expressing themselves online. This may require editing the filters on what is discriminatory language or not, and implementing stringent measures to ensure that those prompting hateful speech and harassment face tangible consequences for their actions.
  2. Incorporating multilingual perspectives. One solution to a lack of diversity in data sets may be obtaining more data and using multilingual sources from other countries. Over 90% of the languages spoken in the world have “little to no support” in language technology, such as in online translation services or as a part of data sets. LLM training data over-represents standard English: 60% of online content is English, despite English being spoken by only 17% of the world. At a minimum, developing translation technologies to include 100% of the online content in LLM training may help AI represent diverse perspectives. 
  3. Balancing training strategies. Platforms such as ChatGPT use a combination of supervised and unsupervised training strategies. Each of these are explained as:
    • Supervised: the LLM is trained to categorize random input values into categories predetermined by humans.
    • Unsupervised: the LLM self-trains by determining patterns in random input values. It determines its own set of categories to organize the data.
      • Unsupervised training is critical because supervised training can be limiting or include human bias in how information should be labelled and organized. However, unsupervised training is difficult because it requires a large data set to determine outliers and draw accurate patterns. Thus, creating a comprehensive data set rooted in solutions 1 and 2 can help make unsupervised training meaningful.

Conclusion

LLMs and generative AI are becoming prevalent sources of information for many netizens globally. However, many technologies have Eurocentric biases that threaten to misrepresent marginalized groups, their vernaculars and their cultures. While it is important to recognize the benefits of LLMs, certain criteria must be met to be used with good intentions. Three proposed solutions are ensuring equal internet access, incorporating multilingual perspectives in the host data and balancing training strategies. Once these three points are achieved, LLMs may actually be vital in the fight against the marginalization of certain groups.

Author