skip to content

For an X.509 leaf certificate, what is the difference between DER and PEM, and which bytes does the signature cover?

level: middleimportance: nice to knowfreq 33%

answer

  1. one value, one encoding
  2. the wrapper is not the object
  3. base64 plus two boundary lines
  4. the body is what gets signed
  5. hash the binary, not the text

basics

~20 s

DER is the single canonical byte encoding of the certificate's ASN.1, and the issuer signs the DER encoding of the TBSCertificate. PEM is base64 armour around those same bytes — a transport wrapper, not a different certificate.

solid answer

~40 s

The Distinguished Encoding Rules give every ASN.1 value exactly **one** byte sequence, which is what makes both a signature and a fingerprint well defined — the more permissive encoding rules DER restricts would allow the same field to be written several ways. The issuer signs the DER encoding of the `TBSCertificate`, so those exact octets are the object. PEM is the same DER, base64-encoded and wrapped between `-----BEGIN CERTIFICATE-----` and `-----END CERTIFICATE-----` lines so it survives being pasted into text; a file may hold several blocks one after another. Converting between the two is lossless and changes nothing signed. A certificate's fingerprint is a hash over the DER of the whole certificate — hashing the armoured text instead gives a digest that identifies nothing.

code

pseudocode · 11 lines
pseudocode
armoured = read_text("leaf-certificate")

body = lines of armoured
       excluding "-----BEGIN CERTIFICATE-----"
       excluding "-----END CERTIFICATE-----"

der = base64_decode(join(body))

fingerprint     = sha256(der)                                  // identifies this certificate
key_digest      = sha256(der_encoding_of(certificate.subjectPublicKeyInfo))  // identifies the key only
identifies_none = sha256(armoured)                             // hashes the wrapper too

go deeper

for a junior

Know that the binary encoding and the base64 armour hold the same certificate, and that converting between them changes nothing about it.

for a middle

Explain why the encoding must be canonical for a signature to be checkable, and say which bytes the issuer actually signed.

for a senior

Anticipate the failures: a re-encoded certificate whose signature no longer verifies, a deduplication key computed over a text wrapper, a key digest mistaken for a certificate digest.

for a principal

Set the rule for how certificates are stored and identified across systems, since a digest that is not over the canonical bytes is not an identifier anyone else can reproduce.

A monitor ingesting certificates from many sources sees the same certificate arrive in two shapes, and has to decide whether they are the same object. The answer turns on what "distinguished" means. ## Why the encoding has to be canonical ASN.1 describes a certificate's *structure*: a sequence with a version, a serial number, names, a key, extensions. Encoding rules turn that structure into bytes, and the general rules are permissive — the same value can legitimately be written in more than one way, for example with different length forms. That flexibility is fatal to a signature. If a verifier re-encoded a parsed certificate and got even one octet different from what the issuer encoded, the hash would differ and a perfectly good signature would fail. The **Distinguished** Encoding Rules remove the choice: one value, one encoding. That single property is what makes the following well defined: - the byte sequence the issuer hashes and signs; - the byte sequence a verifier hashes and checks; - a fingerprint, which is only an identifier if everyone computes it over the same bytes. ## What is actually signed The `signatureValue` covers the **DER encoding of the `tbsCertificate`** — the body alone, not the outer structure, not the algorithm identifier beside it, and obviously not itself. The practical consequence is one a senior engineer meets eventually: a tool that parses a certificate into objects and re-serialises them can produce bytes that differ from the issued ones, and the signature then fails on a certificate that is entirely legitimate. The rule is to keep and verify the original encoded bytes rather than a round-tripped reconstruction. ## PEM is armour, not a format | | DER | PEM | |---|---|---| | What it is | The canonical binary encoding | Base64 of that binary, plus header and footer lines | | What it is for | Signing, verifying, hashing, transmitting | Surviving text channels: config files, message bodies, copy-paste | | Contents | The certificate | The same certificate, re-dressed | | Multiple certificates | One structure per file | Several blocks may be concatenated in one file | | Safe to hash as an identifier | Yes | No — the wrapper is included in the digest | The armour is `-----BEGIN CERTIFICATE-----`, a base64 body conventionally wrapped at 64 characters per line, and `-----END CERTIFICATE-----`. None of it is part of the certificate. Strip the lines, base64-decode the body, and the DER comes back byte for byte; re-armour it and you have the original text. Nothing in the structure, the validity window or the signature has changed, because nothing about the certificate was ever in the wrapper. ## Three different things called "the fingerprint" This is where deduplication goes wrong in a monitor. At least three digests are in circulation and they are not interchangeable: 1. A hash over the **DER of the whole certificate** — the value normally meant by "the certificate's fingerprint". Two certificates with this digest in common are the same certificate. 2. A hash over **only `SubjectPublicKeyInfo`** — identifies the *key*, and is deliberately stable across reissue, because a certificate reissued for the same key carries the same key digest with a new serial number and a new window. 3. A hash over **the armoured text** — identifies nothing. Add a trailing newline, rewrap the lines, prepend a comment, and it changes while the certificate does not. Confusing the first two is the more interesting error: they answer different questions, and a system that stores one while believing it has the other will draw wrong conclusions the first time a certificate is reissued for an unchanged key. ## Practical consequences - Deduplicate on a digest over the DER, never over the file as received. - Keep the original encoded bytes for anything that will be verified later; re-encoding is not a safe round trip. - Expect one file to contain several armoured blocks, and split on the boundary lines rather than assuming one certificate per file. - Treat a `.pem` or `.der` file extension as a hint about packaging and nothing more; the contents decide.

  • A tool parses a certificate, re-encodes it, and the signature stops verifying. Why?
    The signature covers the exact octets the issuer encoded. A re-encode that differs anywhere — a length form, a field the parser normalised or dropped — produces a different hash input, so verification fails on a genuine certificate. Verify the bytes as received, not a reconstruction.
  • The same certificate as a binary file and as an armoured text file give different digests. Which is the fingerprint?
    The digest over the binary DER. The armoured file includes the boundary lines, the base64 alphabet and the line breaks in the hash input, all of which can change without the certificate changing at all.
  • Does converting from DER to armoured text and back change anything a verifier checks?
    No. The decode reproduces the identical octets, so the signed body, the validity window and every extension are unchanged. Only the packaging differed, and packaging is not part of the structure.

saying these in an interview costs you the question

  • Treats DER and PEM as different certificate formats with different contents
  • Computes the fingerprint by hashing the armoured text file
  • Says the signature covers the whole file including the boundary lines
  • Assumes any ASN.1 encoding of the same fields produces the same bytes
  • Uses a certificate digest and a public-key digest interchangeably