Encoding desk
Text to Hex
Encode text as uppercase UTF-8 bytes in a clean spaced hex format.
Text to Hex: Expose the UTF-8 bytes of a text string as spaced, uppercase hexadecimal pairs. Text to Hex makes spaces, line endings and multibyte characters visible. Use it to compare copied text or prepare an encoding example; the byte sequence represents text, not a numeric base conversion. Runs 100% locally in your browser with zero server file uploads.
- Category
- Developer tools
- Runs
- In your browser
- Cost
- Free · no sign-up
- Availability
- Ready to use
Runs entirely in your browser
Result
ReadyOutput is uppercase UTF-8 hex with spaces between bytes. Text is encoded locally and never uploaded.
Byte-rendering rule with a mixed example
First encode the input as UTF-8 bytes b_1 through b_m, then write each byte as two uppercase hex digits with one space between bytes. For m > 0, displayed length = 2m + (m − 1) = 3m − 1. Input Aé has bytes 65, 195 and 169 in decimal, giving 41 C3 A9. With m = 3 the display uses 3×3 − 1 = 8 characters. The A contributes one byte and é contributes two; the space between C3 and A9 is output formatting, not an input space. URL encoding prepares text for URL contexts using its own escaping rules.
UTF-8 validity and capacity
RFC 3629 (https://www.rfc-editor.org/rfc/rfc3629) documents UTF-8 byte sequences. A lone UTF-16 surrogate in an input string is replaced during encoding by U+FFFD, shown as EF BF BD; this is different from the scalar-only character tools, which reject surrogate values. Text is limited to 100,000 code points, with the editor additionally counting UTF-16 units towards its 100,000-character allowance; many emoji consume two of those units. Displayed output is capped at 400,000 characters including separating spaces. Consequently not every maximum-length non-ASCII input fits: multibyte encoding expands the result. Save the source text and try smaller sections when the output limit is reached.
How to use it
- Enter the value in the format shown below.
- Run the conversion and check the result against your source convention.
- Copy or download the output in the form required by your destination.
Privacy & limitations
Conversion runs in your browser. Your values and results are not uploaded.
Related tools
Frequently asked questions
Why does one accented letter produce several byte pairs?
UTF-8 uses one to four bytes per Unicode scalar, so byte count and visible-letter count differ. The letter é produces C3 A9 and an emoji can require four bytes. A composed letter and a letter followed by a combining mark can look alike while yielding different byte sequences.
Are newlines and spaces removed before encoding?
No. Ordinary space becomes 20, a tab becomes 09, LF becomes 0A and CRLF becomes 0D 0A. The converter does not normalise the text or its line endings. Browser pasting may already have changed a source’s line endings, so this shows the bytes of the text actually entered.
Does this output UTF-16, BCD or little-endian integers?
It always encodes UTF-8 text. Digits typed as text become character bytes: 42 gives 34 32, not the ordinary integer hex 2A or packed BCD 42. There is no numeric field width or endianness setting, and reversing these byte pairs would corrupt many multibyte characters.
Free tool · runs in your browser · no account required