Encoded Erasure: Online Name Misrepresentation through the History of Unicode

Bríd-Áine Parnell, 2026

You may have skipped past my name at the start of this blog (we all do), but I invite you to look at it now. It’s likely to be a name you haven’t come across before, because it’s in the Irish language – even in Ireland, it’s a relatively rare name – and little did my parents know when they gave it to me that it would become the inspiration for my PhD research.

When I emigrated to the UK from Ireland in 1999, it made sense to me that no computer systems could accept my name. Instead of typing it as Bríd-Áine, I had to omit my diacritic marks, those accents over the vowels, and often the hyphen too, to fit into bureaucratic boxes. Since I was now living in the birthplace of the English language (and it was perhaps a different time) I adapted, and eventually I came to write my name without its accents even in my own hand.

But as the multilingual internet became more prevalent, as smartphones started to support more and more languages, I started to be able to type my name as it was meant to be. I became retrospectively angered by those initial omissions and much more frustrated with the many continued instances where my name was refused.

An example of online name refusal with an Irish-language name containing diacritics.
An example of online name refusal with an Irish-language name containing diacritics. (Author’s documentation)
A slide showing “Sean”, “Sean”, and “Sean” showing the difference in meaning and pronunciation based on placement of an accent
An example of the differences in meaning and pronunciation based on placement of an accent. (Author’s documentation)

Online Name Misrepresentation

But why, when the Unicode Standard has achieved so much penetration into digital systems, are so many digital identities still reduced to ASCII?

Although Irish is a minority language, this issue does not only occur for minority or digitally-disadvantaged languages. Vast numbers of people have difficulty getting their name accepted in digital forms, when their names are in globe-spanning languages like French, Dutch, and Spanish, dominant European languages such as German, Croatian, or Danish, as well as when their names are in minority languages like Irish or Breton. What all these languages have in common is the use of diacritics or other so-called “special characters” that are refused when digital systems are limited only to the basic Latin script. In other words, when they only accept ASCII (American Standard Code for Information Interchange) encoding.

But why, when the Unicode Standard has achieved so much penetration into digital systems, are so many digital identities still reduced to ASCII? This is the core of the question that brought me, at the kind invitation of the Script Encoding Initiative, first to Berkeley to see the workings of the Unicode Technical Committee and its Script Encoding Working Group, and then to Stanford Libraries with its trove of data on the formation of the Unicode Standard. I wanted to establish the historical context for the encoding of multiple languages in digital systems. How was this accomplished? What were the competing strategies, designs, and motivations for creating the digital text stack we have today? 

Early Unicode Consortium members at work sitting at a large conference table
Early Unicode Consortium members at work (Source: Stanford Libraries Unicode Collection, 1992)

I also hoped to get a sense of that moment in time when a design that essentially ripped up the rulebook was able to gain purchase. Infrastructural theory within science and technology studies (STS) suggests that large technical infrastructures become established and stabilise and once they do, transitions are difficult to achieve. Under Multi-Level Transition (MLT) theory, for example, there must be an evolution through one or more of three levels of interdependent activity for a transition to occur: a technological development, a change in the institutions, norms and practices, or a broader societal trend. STS more broadly asks how societies and their technosciences shape each other within a given historical moment, understanding this relationship as a practice of co-production in which they mutually produce each other. In that moment then, in the late 1980s to 1990s, there was an intervention in the existing system of digital text production. What were the forces that made it possible? 

I also hoped to get a sense of that moment in time when a design that essentially ripped up the rulebook was able to gain purchase.

New scripts and new encodings

During my first few weeks as a visiting scholar, SEI Director Anushah Hossain helped fill my calendar with opportunities to better understand the script encoding process. I attended classes given by Associate Professor Adam Benkato on “Writing Systems of the World” and discussed encoding with the wider team, including Program Manager Helena Kansa and Research Associate Julián Vargo. But perhaps the highlight of this initial period was a talk by invited guests Ibrahima and Abdoulaye Barry on the creation and eventual encoding of the Adlam script for the Fulani language of West Africa.1

Anushah’s work had already introduced me to the concept of neographies – newly invented scripts that are often applied to languages of formerly colonised regions.2 In these cases, a native spoken language was first written down or first widely disseminated in scripts from outside that region, the Latin or sometimes the Arabic script. At times, this meant disagreement in how the language should be written because of a failure of these scripts to accurately represent all the sounds and words of the indigenous language. The Barry brothers set out to first correct this misrepresentation for the Fulani language, spoken by more than 40 million people across West Africa, and then to get their new script encoded so that speakers could communicate digitally. 

A portion of the Adlam character code table
A portion of the Adlam character code table (Source: The Unicode Standard, Version 17.0)

This story resonated strongly with my own work on the effects of the linguistic domination of the Latin alphabet, and particularly the English language. It also spoke to the necessity of digital representation. As the Barry brothers and Anushah both pointed out at the event, for new scripts, legitimacy and spreading usage hinges on the question of whether speakers can use it on their phone.

Watching Unicode at work

In the next stage of my visit, I was able to observe the work of the Script Encoding Working Group (SEWG) and the quarterly meeting of the Unicode Technical Committee (UTC). I was also able to seek feedback – and pique interest in my own project! – by presenting the findings from a survey of users with Irish-language names and their encounters with online systems and digital name misrepresentation. This led to some fantastic conversations in the room, as well as connecting me with many members who have generously given their time in interviews about their experiences past and present.

Through these encounters, the design principles and complicating factors of the ongoing process of standard-making started to emerge. While technical standards often seek to be just that, an engineering solution to a technological problem, they cannot escape their entanglement with the social and even the political world around them. For Unicode, there was a strictly technical question of how the world’s scripts can be encoded in the most technologically useful way possible, but around and underneath this lie the messy complications of human languages and the power and control inherent in how a language is (allowed to be) used.

Bríd-Áine Parnell presenting a lunch talk during the Spring 2026 Unicode Technical Committee meeting
Presenting a lunch talk during the Spring 2026 Unicode Technical Committee meeting

The process of standardisation can often freeze a moment in time, rendering fixed and immutable what was once more malleable and flexible. With Adlam, for example, the Barry brothers spoke about how as the script was taught more and more, educators fed back to them that some of the glyphs for characters were too similar, making it more difficult to learn. But we don’t need to go to a neography to see that language tends to be a living thing, constantly evolving and modified by contact with new speakers, other languages, and even pop culture. Every year, dictionaries from Oxford and Cambridge to Merriam-Webster and Collins choose a “Word of the Year” in English from lists that encompass phrases and words with new meanings (ragebaiting, vibe coding), portmanteaus (situationship, doomscrolling), slang terms (six-seven, aura farming) and newly borrowed words from other languages (lepak [from Malay, meaning to do nothing in particular], affogato [from Italian, a dessert of ice cream with a hot espresso poured on top]).

Text encoding then has to find a way to freeze certain elements of language, such as the individual characters of a script, and still leave enough flexibility to allow language to evolve. But as the Adlam example shows, even the characters of a script – or at least, their expression in glyphs – are not as stable as the Unicode Standard wishes or imagines them to be.

In the archives

Not only does standardisation tend to capture a moment in time, it is also very much of its time. The Unicode-encoded universe continues to grow, both in terms of the scripts and languages it enables and in the associated tools, libraries, and software developed by the Unicode Consortium for localisation and globalisation. But the standard’s early evolution locked in place a set of design principles that remain strongly in influence today. These principles were first proposed by Joseph Becker in his famous Scientific American article in the late 1980s, but were subsequently adapted and modified to meet the requirements of getting Unicode done. 

A prominent lamp featured on a desk in the Stanford Archive Library.
Working hard in the archives (Author’s documentation)
Boxes of primary materials and documents in the Stanford Archive Library.
Working hard in the archives (Author’s documentation)
Notes and check-our permissions slip in the Stanford Archive Library.
Working hard in the archives (Author’s documentation)

The archive documents preserved and donated by Unicode founding member Ken Whistler tell a story of aspiration and compromise, influenced by the sociopolitical world that Unicode was a part of. Two major necessities drove these alterations in the original principles. First, the early Unicoders needed to provide existing text processing firms and implementers with an on-ramp – an easy way to adopt Unicode without throwing their existing products out. Second, Unicode had to find a way to work with the International Organisation for Standardisation (ISO), who were also trying to update text encoding to facilitate more languages.

I still see this tension in the work of Unicode today. There is a commitment to holding to the key principles of non-duplication of characters and encoding once and forever, yet adhering to these principles is so much more complicated and socially entangled than it first appears.

Conclusion

My time with the SEI has been foundational in my research, answering some of my questions about the digital text stack and throwing up so many more. I am still processing the wide range of data that I was able to gather from the archives, from my interviews with Unicode members, both old and new, and from my observations of their work. 

However, this history provides an important context for the questions of today. With so much work done to make Unicode a success, why do digital identities remain a stubborn outlier? The answer may lie in what standards hold in place and what stays flexible. Standards do not operate in a blank space. There are already strictures, habits, and existing technologies in the sociotechnical space that standards enter into, as it was when Unicode was being developed. Some of them disappear into the technological infrastructure and only become clear when they become a problem. Others change and evolve along their own path, surfacing new challenges for existing standards. 

This is a moment in time in which de facto standards for digital identities may be set. Governments around the world are developing digital identity systems for citizen access to services and passports increasingly include a “chip” that holds a digital version of the printed information. An intervention that highlights the vast numbers of people misnamed by current digital identities has the potential for real impact, motivating designers to create systems that allow diacritics and celebrate, instead of erasing, cultural representation.


Bríd-Áine Parnell is a PhD candidate on the Designing Responsible Natural Language Processing at the University of Edinburgh’s School of Informatics. Her work on personal names and digital identity investigates the social, political and technological entanglements of storing and processing personal data in many-layered information infrastructures, where the capabilities of the cutting edge are often limited by the affordances at the foundations. She has a background as a technology journalist and in interdisciplinary studies including the study of sociotechnical systems, data capitalism and colonialism, media and communication studies, sociolinguistics, computer science, and algorithmic justice. Learn more about her work here.
Headshot of SEI visiting researcher Bríd-Áine Parnell
  1. The Barry brothers spoke at the UCB event co-organized by SEI, “Inventing Adlam: The Origins and Impact of a New African Writing System“. Watch the full recording here
    ↩︎
  2. The article offers a definition of neographies and describes the dilemmas involved in their encodings. “Unsettled Scripts: Script Invention and the Bind of Digital Recognition” is available here.
    ↩︎