Text / Encoding

Byte Counter (UTF-8, UTF-16, and SMS)

Count a text's bytes in UTF-8 and UTF-16, its code points and visible characters, and the SMS parts it needs, and cut it to a byte limit without breaking a character.

Byte Counter (UTF-8, UTF-16, and SMS): Different systems measure length differently. UTF-8, used by web pages and most databases, takes 1 byte for ASCII, 2 for most accented and Greek or Cyrillic letters, 3 for Chinese, Japanese, and Korean, and 4 for emoji. JavaScript counts UTF-16 code units, Python counts code points, and people see grapheme clusters, so a skin-toned emoji is one character to a reader but two code points and eight UTF-8 bytes. SMS messages fit 160 characters from the GSM alphabet, or 70 if any other character appears. Runs 100% locally in your browser with zero server file uploads.

Category
Text tools
Runs
In your browser
Cost
Free · no sign-up
Availability
Ready to use
Byte counterLocal processing

Runs entirely in your browser

UTF-8 bytes24web pages, files, most databases
UTF-16 bytes26JavaScript, Java, Windows
Characters you see10grapheme clusters
Code points11what Python's len() counts
JavaScript length13UTF-16 code units
Words and lines3 · 1
SMS parts1UCS-2 · 13 of 70 per part

These characters are outside the GSM alphabet, so the whole message is sent as UCS-2 with 70 characters per part: 👍 🏽 日 本 語

Contains 4-byte characters such as emoji: MySQL needs the utf8mb4 character set to store them; its older utf8 (utf8mb3) set rejects them.

Limits measured in bytes

Many limits count bytes, not characters: database columns and index keys, HTTP headers and cookies (about 4 KB per cookie), file names (255 bytes on most file systems), and payment or messaging APIs. A limit that fits 255 English letters may hold only 85 Chinese characters or 63 emoji.

For limits counted in characters, such as social posts, use the character counter; to find out how a file is encoded, the text encoding detector.

Keeping SMS messages short

Replace curly quotes and apostrophes with straight ones, dashes with hyphens, and avoid emoji, so the message stays in the GSM alphabet and fits 160 characters per part instead of 70. Each extra part is usually billed as a separate message.

How to use it

  1. Type or paste the text.
  2. Read the byte counts, character counts, and SMS parts.
  3. Enter a byte limit to see whether the text fits, and get a safely cut version.

Privacy & limitations

The text is measured in your browser and never uploaded.

Related tools

Frequently asked questions

Why does my database say the text is too long?

Column limits may count bytes rather than characters: MySQL index prefixes and many older systems use bytes, so 255 Chinese characters need 765 bytes in UTF-8.

Why does one emoji send my SMS as Unicode?

SMS uses the 7-bit GSM alphabet when every character is in it; one character outside it, such as an emoji or a curly quote, switches the whole message to UCS-2, which fits 70 characters per part instead of 160.

What is a long SMS split into?

Parts of 153 GSM characters or 67 UCS-2 units, since each part carries a header to join them on the phone; some characters, such as € and [ ], count twice in GSM.

Free tool · runs in your browser · no account required