In OpenCV, how do you get each blob's area and centroid from a binary mask?
answer
- m00 is the area
- divide the first-order moments by it
- label zero is not an object
- one route counts pixels, the other integrates a polygon
- guard against a zero denominator
basics
~20 sTwo routes. Per contour: cv2.moments gives m00 as the area and m10/m00, m01/m00 as the centroid. Per labelled region: cv2.connectedComponentsWithStats returns stats with CC_STAT_AREA in pixels plus a centroids array - remember label 0 is the background.
solid answer
~40 sWith contours, `m = cv2.moments(cnt)` returns a dict of raw, central and normalized moments; `m["m00"]` is the area and the centroid is `(m["m10"]/m["m00"], m["m01"]/m["m00"])`. Guard `m00 == 0` - a degenerate or collinear contour divides by zero. With labeling, `retval, labels, stats, centroids = cv2.connectedComponentsWithStats(binary, connectivity=8)` hands you everything at once: `stats[i, cv2.CC_STAT_AREA]` is the exact foreground **pixel count**, `stats[i]` also carries `CC_STAT_LEFT/TOP/WIDTH/HEIGHT` for a bounding box, and `centroids[i]` is a subpixel centre. Label 0 is the background, so the blob count is `retval - 1` and you iterate from 1. The two areas are not identical: `m00` integrates the contour polygon and therefore includes any holes, while `CC_STAT_AREA` counts only true foreground pixels. Pick by which definition your specification means.
code
python · 16 linesimport cv2
import numpy as np
img = np.zeros((200, 200), np.uint8)
cv2.rectangle(img, (20, 20), (60, 60), 255, -1)
cv2.circle(img, (140, 140), 20, 255, -1)
n, labels, stats, centroids = cv2.connectedComponentsWithStats(img, connectivity=8)
print("blobs:", n - 1) # label 0 is the background
for i in range(1, n):
print(stats[i, cv2.CC_STAT_AREA], centroids[i])
contours, _ = cv2.findContours(img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
m = cv2.moments(contours[0])
if m["m00"] != 0:
print(m["m00"], m["m10"] / m["m00"], m["m01"] / m["m00"])go deeper
Know the centroid formula from cv2.moments - cx is m10 over m00, cy is m01 over m00 - and that m00 is the area. Being able to name cv2.contourArea and cv2.boundingRect rounds it out.
Explain both routes and their return shapes, the CC_STAT_ column constants, and why the object count is retval minus one. Name the m00 zero guard as an actual failure you code around.
Show that you know the two area definitions differ - polygon integral including holes versus a literal pixel count - and choose deliberately based on what the inspection spec means, plus the 4-versus-8 connectivity call on touching parts.
Own the measurement contract: which definition of area the acceptance criteria encode, how pixel-to-millimetre calibration and boundary discretization bound the achievable precision, and when quantisation error makes a classical threshold-and-count approach unfit for the tolerance being claimed.
## The two APIs and what each is really for After thresholding and morphological clean-up you have a binary mask, and the next question is always the same: how many objects, how big, and where. OpenCV offers two paths. **Contours** (`cv2.findContours` + `cv2.moments` and friends) give you an ordered boundary curve, which unlocks shape: perimeter, convexity, polygon approximation, rotated bounding boxes, shape matching. If you need to know that something is a hexagon, you need contours. **Connected-component labeling** (`cv2.connectedComponentsWithStats`) gives you a label image plus a statistics table in a single pass. It answers count, pixel area, bounding box and centroid directly, and it hands back `labels`, an integer image you can compare against to build a per-blob mask instantly. If you need to know how many blobs there are and how big each is, this is the shorter and usually faster road. ## Moments in practice `cv2.moments(array, binaryImage=False)` accepts either a contour or a single-channel image and returns a dict of 24 values. The names follow a convention: `m00`, `m10`, `m01`, `m20`, ... are raw moments; `mu20`, `mu11`, `mu02`, ... are central moments (translation invariant); `nu20`, `nu11`, ... are normalized central moments (also scale invariant). The two you use daily: - `m00` is the area. For a contour, it equals `cv2.contourArea` of the same curve. - The centroid is `cx = m["m10"] / m["m00"]`, `cy = m["m01"] / m["m00"]`. Intuitively, a raw moment is a coordinate-weighted sum; dividing by the total weight gives the mean position. The mandatory guard is `m00 == 0`. It happens more often than people expect: a contour of a single pixel, a perfectly collinear run of pixels, or a degenerate curve produced by aggressive approximation. In Python that is a `ZeroDivisionError` in the middle of a frame loop, so filter contours by `cv2.contourArea` before measuring, or test `m00` explicitly. When you pass an *image* rather than a contour, `binaryImage=True` makes every non-zero pixel count as 1 instead of contributing its intensity - that is the flag to set for masks, and forgetting it on a 0/255 mask scales the moments by 255 (the centroid survives, the area does not). The higher moments are not decoration. `cv2.HuMoments(m)` derives seven values invariant to translation, scale and rotation, which `cv2.matchShapes` uses to compare two contours regardless of pose - the classical answer to "is this the same part." ## Connected components in practice ``` n, labels, stats, centroids = cv2.connectedComponentsWithStats(binary, connectivity=8) ``` `n` is the number of labels **including the background**, so the object count is `n - 1`. `labels` is an `int32` image where each pixel holds its component's label. `stats` has shape `(n, 5)`; use the named column constants rather than integers: `cv2.CC_STAT_LEFT`, `CC_STAT_TOP`, `CC_STAT_WIDTH`, `CC_STAT_HEIGHT`, `CC_STAT_AREA`. `centroids` has shape `(n, 2)` in float64, so centres are subpixel. **Row 0 is the background**, and it is the single most common bug on this API: iterate from 1, and if you sort blobs by area without dropping row 0 the background wins every time, because it is usually the largest region in the image. `connectivity` is 8 by default (diagonal neighbours join) or 4 (only orthogonal). This is a real decision, not a formality: two blobs touching only at a corner are one component under 8-connectivity and two under 4-connectivity. On thin diagonal structures, 4-connectivity fragments them. Building a mask for one blob is then a one-liner: `mask = (labels == i).astype(np.uint8) * 255`, which is far cheaper and more exact than redrawing a contour. ## Why the two areas disagree This distinction is worth stating precisely because it decides pass/fail in real inspection systems. `m00` and `cv2.contourArea` integrate the **polygon** traced by the contour, via Green's theorem. Consequences: the value includes any holes inside the shape, and because the polygon runs through boundary pixel centres it differs from a raw pixel count by roughly a half-pixel band around the perimeter - a few percent on small blobs. `stats[i, cv2.CC_STAT_AREA]` is a literal count of pixels carrying that label. Holes are excluded automatically, since hole pixels belong to a different component. So for a washer, the contour route reports the full disc and the labeling route reports the material. Neither is wrong; they answer different questions. If your spec says "surface area of material," take the pixel count. If it says "footprint," take the polygon. ## Bounding geometry `cv2.boundingRect(cnt)` returns `(x, y, w, h)`, an axis-aligned box - the same information as the four `CC_STAT_*` columns. `cv2.minAreaRect(cnt)` returns a **rotated** rectangle as `((cx, cy), (w, h), angle)`, which is what you want for elongated parts at arbitrary orientation; `cv2.boxPoints(rect)` converts it to four corner points for drawing. The ratio of `contourArea` to the min-area-rect area is a cheap, orientation-independent fill ratio for rejecting badly-shaped blobs. ## Choosing If the pipeline is count-measure-locate, reach for `connectedComponentsWithStats`: one call, exact pixel areas, no zero-area edge case, and a reusable label image. If the pipeline needs shape - polygon approximation, convexity, rotated boxes, shape matching, hole structure - use contours, and filter by area before you compute anything that divides.
- cv2.connectedComponentsWithStats returns 5 for a mask you believe holds four blobs. Is that a bug?No. The returned count includes the background as label 0, so five labels means four objects: the blob count is retval - 1. The matching discipline is to iterate stats and centroids from index 1, and never to take an argmax over the area column without excluding row 0, since the background is usually the largest region.
- Your centroid computation raises ZeroDivisionError mid-stream. Why, and how do you prevent it?Some contour has m00 equal to zero - a single-pixel blob, a collinear run, or a curve degenerated by polygon approximation. Guard it directly by skipping contours whose m00 is zero, and more usefully filter the contour list by cv2.contourArea against a minimum size before measuring anything, which also removes noise you did not want to report.
- When would you compute the area with contourArea rather than CC_STAT_AREA?When the specification means the footprint of the outer boundary rather than the material - for example the swept envelope of a part with cut-outs - or when you already have the contour for shape analysis and want a consistent basis with its perimeter and rotated bounding box. Just remember contourArea includes holes and comes from a polygon integral, not a pixel count.
- Two parts touch at a single diagonal pixel. How does connectivity change the result?With connectivity=8, the default, diagonal neighbours are joined and the two parts are reported as one component with a merged area and a centroid between them. With connectivity=4 only orthogonal neighbours join, so they separate correctly - at the cost of fragmenting genuinely thin diagonal structures elsewhere in the frame.
saying these in an interview costs you the question
- Counts blobs as retval instead of retval minus one
- Forgets to guard m00 equal to zero
- Treats contourArea and pixel count as identical
- Reads centroids row 0 as the first object
- Ignores connectivity when blobs touch diagonally