LLVM and Clang use a subset of the Unicode Character Database (UCD) for identifiers, \N{...} named characters, diagnostic column width, and simple case folding. The corresponding tables are generated from the UCD into Clang and LLVM Support.
Download the following files:
https://www.unicode.org/Public/UCD/latest/ucdxml/ucd.nounihan.flat.zip https://www.unicode.org/Public/UCD/latest/ucd/UnicodeData.txt https://www.unicode.org/Public/UCD/latest/ucd/NameAliases.txt https://www.unicode.org/Public/UCD/latest/ucd/extracted/DerivedName.txt
Unzip ucd.nounihan.flat.zip to get ucd.nounihan.flat.xml.
Build the generators from an LLVM build directory that has utilities enabled (LLVM_BUILD_UTILS, the default) and libxml2 (LLVM_ENABLE_LIBXML2, also the default):
ninja -C <build> UnicodeCharSetsGenerator UnicodeNameMappingGenerator
Then run them from the llvm-project root as shown below.
UnicodeCharSetsGenerator writes the identifier character sets used by the lexer (XID_Start, XID_Continue, and the mathematical compatibility notation profile) and the printable / formatting / combining / East-Asian-width sets used by LLVM Support.
<build>/bin/UnicodeCharSetsGenerator ucd.nounihan.flat.xml \ clang/lib/Lex/UnicodeCharSetsGenerated.cpp \ llvm/lib/Support/UnicodeCharSetsGenerated.cpp clang-format -i clang/lib/Lex/UnicodeCharSetsGenerated.cpp \ llvm/lib/Support/UnicodeCharSetsGenerated.cpp
The C99 and C11 identifier tables and the whitespace table in clang/lib/Lex/UnicodeCharSets.h are maintained by hand.
UnicodeNameMappingGenerator writes the name-to-codepoint trie used by \N{...} and llvm::sys::unicode::nameToCodepointStrict / nameToCodepointLooseMatching.
<build>/bin/UnicodeNameMappingGenerator UnicodeData.txt NameAliases.txt \ llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp clang-format -i llvm/lib/Support/UnicodeNameToCodepointGenerated.cpp
Algorithmically derived names (CJK UNIFIED IDEOGRAPH-*, and so on) are not in those files. Update GeneratedNamesDataTable in llvm/lib/Support/UnicodeNameToCodepoint.cpp from DerivedName.txt.
llvm/utils/unicode-case-fold.py fetches CaseFolding.txt and writes llvm/lib/Support/UnicodeCaseFold.cpp:
llvm/utils/unicode-case-fold.py \ https://www.unicode.org/Public/UCD/latest/ucd/CaseFolding.txt \ > llvm/lib/Support/UnicodeCaseFold.cpp
After regenerating the tables, update llvm/unittests/Support/UnicodeTest.cpp and clang/test/Lexer/unicode.c for new characters and changed name ranges.