skip to content

Which image data does cv2.imread() discard by default, and how do you keep it?

level: middleimportance: should knowfreq 44%

answer

  1. the default flag normalises every file the same way
  2. three channels, eight bits, always
  3. transparency and precision go missing quietly
  4. IMREAD_UNCHANGED returns it as stored
  5. check img.shape and img.dtype after reading

basics

~10 s

The default flag cv2.IMREAD_COLOR always yields a 3-channel 8-bit BGR array: it drops any alpha channel and truncates 16-bit files to 8 bits. cv2.IMREAD_UNCHANGED returns the data as stored, alpha and bit depth included.

solid answer

~40 s

`cv2.imread(path)` uses `cv2.IMREAD_COLOR`, which normalises whatever was in the file into three 8-bit BGR channels. A PNG with transparency loses its alpha channel; a 16-bit depth map or medical scan is silently scaled down to 8 bits, destroying precision; a grayscale file is expanded into three identical channels. None of that raises. `cv2.IMREAD_UNCHANGED` returns the decoded data as stored — four channels (BGRA) when alpha is present, `uint16` dtype for 16-bit files — but the docs also note that it ignores EXIF orientation, so a phone photo can come back rotated relative to the default read. `cv2.IMREAD_GRAYSCALE` gives one channel, and `cv2.IMREAD_ANYDEPTH` preserves depth without preserving alpha. After any read, `img.shape` and `img.dtype` tell you exactly what you got — check them rather than assuming.

code

python · 10 lines
python
import cv2

default = cv2.imread("logo_16bit_rgba.png")
print(default.shape, default.dtype)   # (H, W, 3) uint8  - alpha and depth gone

exact = cv2.imread("logo_16bit_rgba.png", cv2.IMREAD_UNCHANGED)
print(exact.shape, exact.dtype)       # (H, W, 4) uint16 - as stored

cv2.imwrite("out.png", exact)                                   # keeps alpha + depth
cv2.imwrite("out.jpg", exact, [cv2.IMWRITE_JPEG_QUALITY, 95])   # loses both

go deeper

for a junior

Know that the plain cv2.imread call always gives you three 8-bit BGR channels, and that cv2.IMREAD_UNCHANGED is the flag to reach for when a PNG has transparency you need to keep.

for a middle

Be able to name all three normalisations the default performs — alpha dropped, depth truncated, grayscale expanded — and to prove which one you got by printing img.shape and img.dtype. Know that imwrite picks the format from the extension.

for a senior

Show judgment about when fidelity matters: depth maps, masks and scientific imagery need IMREAD_UNCHANGED plus a format that can hold the result, and a JPEG round-trip silently undoes it. Talk about asserting dtype at ingest so a precision loss fails loudly.

for a principal

Own the storage-format policy across a pipeline — which stage is allowed to be lossy, what the archival format is, and how bit depth and alpha are carried end to end so that no downstream consumer has to guess what a file contains.

## The default is a normaliser, not a faithful reader The convenience of `cv2.imread(path)` is that whatever the file contains, you get the same thing back: an `(H, W, 3)` array of `uint8` in BGR order. That uniformity is why the default exists, and it is also what makes it lossy. Three separate normalisations happen, all silently. **Alpha is dropped.** A PNG or TIFF with a transparency channel decodes to three channels. There is no warning; the array simply has no fourth channel, and anything that depended on transparency — compositing a logo, masking a cut-out — quietly stops working. **Bit depth is reduced.** 16-bit PNG and TIFF files are common in depth sensing, satellite imagery, microscopy and medical imaging. The default read converts them to 8 bits, collapsing 65,536 levels into 256. The image still *looks* fine; what has vanished is the precision your downstream analysis needed. This is the most damaging of the three, because the result is plausible rather than obviously broken. **Channel count is expanded.** A single-channel grayscale file becomes three identical channels. That wastes memory and hides the fact that the source had no colour information at all. ## The flags - `cv2.IMREAD_COLOR` (default) — 3-channel 8-bit BGR, always. - `cv2.IMREAD_UNCHANGED` — returns the image as stored: 4 channels for BGRA sources, `uint16` for 16-bit files, 1 channel for grayscale sources. The documentation also notes that this flag **ignores EXIF orientation**, which matters for phone photos: the default read applies the EXIF rotation, so switching to `IMREAD_UNCHANGED` can change the image's orientation as a side effect you did not ask for. - `cv2.IMREAD_GRAYSCALE` — one 8-bit channel, converted by the decoder. Note this is not always bit-identical to reading colour and then calling `cvtColor(..., COLOR_BGR2GRAY)`, because the conversion happens at a different point in the decode pipeline. - `cv2.IMREAD_ANYDEPTH` — preserves 16-bit or 32-bit depth. Combine it with `cv2.IMREAD_ANYCOLOR` to preserve both depth and the stored channel layout. - `cv2.IMREAD_REDUCED_COLOR_2`, `_4`, `_8` — decode at half, quarter or eighth resolution. Genuinely useful for thumbnails, because the decoder does less work rather than decoding fully and then resizing. - `cv2.IMREAD_IGNORE_ORIENTATION` — read the pixels as stored without applying the EXIF rotation. ## Verify what you got Two attributes settle every question: ``` img = cv2.imread(path, cv2.IMREAD_UNCHANGED) print(img.shape) # (H, W) or (H, W, 3) or (H, W, 4) print(img.dtype) # uint8 or uint16 ``` A `(H, W, 4)` array will surprise code written for three channels — indexing `img[:, :, 2]` still gives red, but anything that assumes a three-channel layout, or that a `uint8` range is 0-255, is now on thin ice. A `uint16` array in particular breaks assumptions everywhere downstream: `cv2.imshow` expects to scale it differently, arithmetic saturates at a different ceiling, and models expecting 0-255 inputs receive values up to 65535. ## The writing side `cv2.imwrite` chooses the file format from the **filename extension** — there is no format parameter. That makes the extension a functional decision, not cosmetic: - JPEG is lossy, 8-bit, and has no alpha channel, so writing a BGRA or `uint16` array to `.jpg` discards exactly what `IMREAD_UNCHANGED` worked to preserve. - PNG is lossless and supports alpha and 16-bit data. - TIFF handles high bit depths; OpenEXR and HDR handle 32-bit float. Encoder options are passed as a flat list of key/value integers, for example `cv2.imwrite(path, img, [cv2.IMWRITE_JPEG_QUALITY, 95])` or `[cv2.IMWRITE_PNG_COMPRESSION, 9]`. And `imwrite` expects **BGR**: handing it an array you converted to RGB for a model writes a colour-swapped file. ## The practical rule Default-read when the image is ordinary photographic input to a vision pipeline that wants 8-bit BGR anyway — which is most of the time, and is why the default is what it is. Reach for `IMREAD_UNCHANGED` the moment transparency or bit depth carries information: masks, depth maps, scientific and medical imagery, and any round-trip where the file you write must preserve what the file you read contained. Then assert on `shape` and `dtype` rather than trusting the file extension to tell you what is inside.

  • You switch a phone-photo pipeline to IMREAD_UNCHANGED and images start coming out rotated. Why?
    The default read applies the EXIF orientation tag, while IMREAD_UNCHANGED is documented to ignore it, so the pixels arrive in their stored orientation. Either rotate explicitly based on the EXIF tag, or keep the default read if you only needed the raw depth or alpha for part of the pipeline. cv2.IMREAD_IGNORE_ORIENTATION makes the same choice explicit for other flags.
  • How does cv2.imwrite decide which format to write?
    Purely from the filename extension — there is no format argument. So '.jpg' means lossy 8-bit with no alpha, and '.png' means lossless with alpha and 16-bit support. Encoder settings go in a flat list of key/value ints, such as [cv2.IMWRITE_JPEG_QUALITY, 95]. imwrite also expects BGR, so writing an RGB array produces a colour-swapped file.
  • What goes wrong if a 16-bit depth image is read with the default flag?
    It is scaled down to uint8, so 65,536 levels collapse into 256. Nothing raises and the image still renders plausibly, but the quantisation destroys the fine distance differences the depth map existed to record. Read it with cv2.IMREAD_UNCHANGED or cv2.IMREAD_ANYDEPTH and check that dtype is uint16.

saying these in an interview costs you the question

  • Assuming imread preserves a PNG's alpha channel by default
  • Thinking a 16-bit file stays 16-bit without a flag
  • Believing cv2.imwrite takes a format argument
  • Writing a BGRA or uint16 array to .jpg and expecting fidelity
  • Trusting the file extension instead of checking shape and dtype

context