File tool / 05
Text Encoding Detector
Check BOM markers, ASCII compatibility, and strict UTF-8 validity in a local file sample.
Text Encoding Detector: TOEA detects UTF-8 and UTF-16 byte-order marks, pure ASCII, and strictly valid UTF-8; other bytes receive a cautious unknown/legacy result. Runs 100% locally in your browser with zero server file uploads.
- Category
- Developer tools
- Runs
- In your browser
- Cost
- Free · no sign-up
- Availability
- Ready to use
Runs entirely in your browser
When the result is legacy or unknown
Text that is not valid UTF-8 is most often in an older single-byte encoding. For Western European text that usually means Windows-1252 or ISO-8859-1. Convert it once and keep UTF-8 from then on: iconv -f WINDOWS-1252 -t UTF-8 in.txt > out.txt on Linux or macOS, or open it in an editor that lets you choose the encoding and save as UTF-8. Only the first 1 MB is inspected, so a large file whose first megabyte is plain ASCII can still contain non-ASCII text further on.
Byte-order marks and UTF-16
A UTF-8 byte-order mark (EF BB BF) helps Excel open a CSV with accents correctly, but it breaks other things: a shell script's #! line, PHP output sent before headers, and the first column name of a CSV read by code that does not strip it. UTF-16 files usually come from Windows, for example from Windows PowerShell 5.1's > redirect. Many Unix tools read them as text full of null bytes.
How to use it
- Choose a text-like file up to 50 MB.
- Inspect the first 1 MB locally.
- Review the encoding and confidence label.
Privacy & limitations
The file stays local. Encoding detection without a BOM is heuristic and cannot always distinguish legacy character sets.
Related tools
Frequently asked questions
Why is ASCII called UTF-8 compatible?
ASCII bytes use the same values in UTF-8, so valid ASCII also decodes as UTF-8.
Can it identify every legacy encoding?
No. Many single-byte encodings are ambiguous without language context.
Free tool · runs in your browser · no account required