Skip to content
Quaestio

Base64

Type or paste text and it is encoded or decoded as you go.

0 characters · 0 bytes

Result

Type something in the field above.
Embed

Embed this tool on your site

Copy the code and paste it where the tool should appear, such as a blog post or a school page. It is free, the box has no ads and the tool calculates in the visitor’s browser.

How it works

Base64 rewrites arbitrary bytes using 64 characters that survive transport through systems built for text, such as email and web addresses.

Three bytes become four characters, which makes the result about a third longer than the original. The equals signs at the end are padding, used when the byte count is not divisible by three.

Base64 is encoding, not encryption. Anyone can decode it, so it protects nothing. Use it to move data, never to hide it.

The text is encoded as UTF-8 before it is converted, which is why accented letters and emoji work. The browser’s built-in function only handles characters below 256 on its own.

The URL-safe variant replaces plus with hyphen and slash with underscore. Decoding here accepts either variant, with or without padding.

How the tool works

When encoding, the tool first turns the text into UTF-8 bytes and then has the browser’s btoa function write them as Base64. When decoding, it removes all white space, including line breaks inside the text, turns hyphens and underscores back into plus and slash, and adds equals signs until the length divides by four. The browser’s atob then decodes, and the bytes are read as UTF-8 in strict mode: bytes that do not form valid UTF-8 produce an error message rather than replacement characters. Under the box you see both characters and bytes. Characters are counted as a reader sees them, so an emoji with a skin tone is one character but eight bytes in UTF-8, and it is the bytes that set the length of the result.

Worked example

Hé becomes three bytes in UTF-8: 48 for H and C3 A9 for é. The 24 bits split into four groups of six: 010010, 001100, 001110 and 101001, which are 18, 12, 14 and 41. In the Base64 alphabet, where A–Z are 0–25 and a–z are 26–51, those are S, M, O and p. The result is SMOp, with no padding because three bytes fit exactly into four characters.

Hey! is four bytes and becomes SGV5IQ==. The fourth byte fills only two characters of a new group of four, and the two equals signs pad out the group.

Edge cases

The tool works with text. Base64 that holds binary data, such as an image or a compressed file, will as a rule give the error message when decoded, because the bytes are not valid UTF-8.

Base64 in email is broken into lines: RFC 2045 allows at most 76 characters per line and says line breaks are to be ignored when decoding. Since the tool removes all white space first, an encoded text part from an email can be pasted in as it is.

A string whose length leaves a remainder of 1 when divided by four can never be valid. A lone final character carries only six bits, not enough for a whole byte, so the algorithm in the WHATWG Infra Standard, which atob follows, rejects such input. SGV5I, which is SGV5IQ== with the last three characters cut off, gives the error message, while SGV5IQ without equals signs decodes to Hey!.

Sources

How the tools are checked