File tool / 05

Text Encoding Detector

Check BOM markers, ASCII compatibility, and strict UTF-8 validity in a local file sample.

Text Encoding Detector: TOEA detects UTF-8 and UTF-16 byte-order marks, pure ASCII, and strictly valid UTF-8; other bytes receive a cautious unknown/legacy result. Runs 100% locally in your browser with zero server file uploads.

Runs
In your browser
Cost
Free · no sign-up
Availability
Ready to use
Text encoding detectorLocal processing

Runs entirely in your browser

When the result is legacy or unknown

Text that is not valid UTF-8 is most often in an older single-byte encoding. For Western European text that usually means Windows-1252 or ISO-8859-1. Convert it once and keep UTF-8 from then on: iconv -f WINDOWS-1252 -t UTF-8 in.txt > out.txt on Linux or macOS, or open it in an editor that lets you choose the encoding and save as UTF-8. Only the first 1 MB is inspected, so a large file whose first megabyte is plain ASCII can still contain non-ASCII text further on.

Byte-order marks and UTF-16

A UTF-8 byte-order mark (EF BB BF) helps Excel open a CSV with accents correctly, but it breaks other things: a shell script's #! line, PHP output sent before headers, and the first column name of a CSV read by code that does not strip it. UTF-16 files usually come from Windows, for example from Windows PowerShell 5.1's > redirect. Many Unix tools read them as text full of null bytes.

How to use it

  1. Choose a text-like file up to 50 MB.
  2. Inspect the first 1 MB locally.
  3. Review the encoding and confidence label.

Privacy & limitations

The file stays local. Encoding detection without a BOM is heuristic and cannot always distinguish legacy character sets.

Related tools

Frequently asked questions

Why is ASCII called UTF-8 compatible?

ASCII bytes use the same values in UTF-8, so valid ASCII also decodes as UTF-8.

Can it identify every legacy encoding?

No. Many single-byte encodings are ambiguous without language context.

Free tool · runs in your browser · no account required