The Token Tax Index: 28 Languages, Measured
Token counts for the same text in 28 languages under o200k_base. Japanese costs 118 percent more than English using half the characters. Open dataset, CSV and JSON.

Token counts for the same text in 28 languages under o200k_base. Japanese costs 118 percent more than English using half the characters. Open dataset, CSV and JSON.

The same instruction costs 73 percent more tokens in Japanese than English, using less than half the characters. Measured with the real GPT-5.6 tokeniser, with a tool to try it yourself.