In OpenCV, what does detectAndCompute() return for a detector like SIFT or ORB?
answer
- a tuple, not one value
- one row per keypoint, same order
- pt, size, angle, response
- dtype differs by algorithm
- empty image gives None, not an empty array
basics
~20 sdetectAndCompute returns a two-element tuple: a list of cv2.KeyPoint objects and a NumPy descriptor array with one row per keypoint. Row i describes keypoints[i]. When nothing is detected, the descriptor value is None, not an empty array.
solid answer
~50 sEvery OpenCV Feature2D detector exposes `detectAndCompute(image, mask)`, which in Python returns `(keypoints, descriptors)`. `keypoints` is a list of `cv2.KeyPoint`, each carrying `.pt` (the sub-pixel x,y location), `.size` (the diameter of the meaningful neighbourhood), `.angle` (dominant orientation in degrees, or -1 if undefined), `.response` (detector strength), `.octave` (pyramid level) and `.class_id`. `descriptors` is a NumPy array of shape `(len(keypoints), D)`, and the **row order matches the keypoint list index for index** — that pairing is what later lets a `DMatch.queryIdx` be turned back into a coordinate. The dtype and D depend on the algorithm: `cv2.SIFT_create()` gives 128 float32 values per keypoint, `cv2.ORB_create()` gives 32 uint8 bytes. The `mask` argument, usually passed as `None`, restricts detection to a region. If no keypoints survive, Python gets `None` for descriptors — pass that into a matcher and it raises.
code
python · 10 linesimport cv2
import numpy as np
img = np.zeros((200, 200), np.uint8)
cv2.rectangle(img, (30, 30), (110, 110), 255, -1)
cv2.circle(img, (150, 150), 30, 200, -1)
kp, des = cv2.SIFT_create().detectAndCompute(img, None)
print(len(kp), des.shape, des.dtype) # e.g. 24 (24, 128) float32
print(kp[0].pt, kp[0].size, kp[0].angle, kp[0].response)go deeper
Know the unpacking idiom kp, des = detector.detectAndCompute(img, None), that des has one row per keypoint, and that KeyPoint.pt is an (x, y) coordinate.
Explain what size, angle, response and octave encode, why the descriptor dtype differs between SIFT and ORB, and why the index pairing between the two returns must be preserved.
Demonstrate the defensive habits: guard against None descriptors on low-texture frames, budget descriptor memory per image, and use the mask argument rather than post-filtering keypoints by region.
Frame descriptor width and dtype as a system cost — index size, network transfer and matching latency all scale with them — and set the keypoint budget per image accordingly rather than leaving detector defaults in place.
## The Feature2D contract OpenCV models detectors and descriptors with one base interface. Whether you instantiate `cv2.SIFT_create()`, `cv2.ORB_create()`, `cv2.AKAZE_create()` or `cv2.BRISK_create()`, you get an object with the same three methods: - `detect(image, mask)` — find keypoints only. - `compute(image, keypoints)` — describe an existing keypoint list. - `detectAndCompute(image, mask)` — do both in one pass, which is faster because the scale-space pyramid is built once. In Python `detectAndCompute` returns a tuple that is almost always unpacked as `kp, des = detector.detectAndCompute(img, None)`. ## What a keypoint is A `cv2.KeyPoint` is a described *location*, not the description itself. Its fields: - **pt** — an (x, y) float tuple. Note the ordering: x is the column, y is the row, the opposite of NumPy indexing `img[y, x]`. Mixing these up puts every drawn marker in the wrong place on non-square images. - **size** — the diameter of the image region the descriptor summarises. This is where scale lives: a keypoint found high in the pyramid has a larger size. - **angle** — dominant gradient orientation in degrees, clockwise, or -1 when the detector does not assign one. Rotation invariance comes from describing the patch relative to this angle. - **response** — the detector's own strength score, useful for keeping the strongest N. - **octave** — the pyramid layer the keypoint came from. - **class_id** — free for you to use, for example to tag which object a keypoint belongs to. `cv2.drawKeypoints(img, kp, None, flags=cv2.DRAW_MATCHES_FLAGS_DRAW_RICH_KEYPOINTS)` draws circles sized by `size` and a radius line at `angle`, which is a fast sanity check that scale and orientation look sensible. ## What the descriptor array is The descriptor array is a plain NumPy 2D array. Its row count equals `len(keypoints)`; its column count and dtype are fixed by the algorithm: - SIFT: 128 columns, float32 — a gradient-orientation histogram over a 4x4 grid of 8 bins. - ORB: 32 columns, uint8 — 256 binary comparison bits packed into bytes. - AKAZE (default MLDB descriptor): uint8, also a packed binary string. That dtype is not cosmetic. It dictates the distance function the matcher must use, and it dominates memory: ten thousand SIFT descriptors is about 5 MB, ten thousand ORB descriptors about 320 KB. ## The index pairing is load-bearing Because row i of the descriptor array belongs to `keypoints[i]`, matching can work purely on integers. A `cv2.DMatch` carries `queryIdx`, `trainIdx`, `imgIdx` and `distance`; you recover geometry with `kp1[m.queryIdx].pt` and `kp2[m.trainIdx].pt`. Anything that reorders or filters one list without the other — sorting keypoints by response but not the descriptor rows — silently corrupts every downstream match. Filter with a boolean index applied to both, or filter the matches instead. ## Common surprises **Descriptors can be None.** On a blank, blurred or extremely low-texture image no keypoints survive, and Python receives `None` rather than a `(0, 128)` array. `bf.match(des1, des2)` then raises. A `if des1 is None or des2 is None: continue` guard belongs in any batch pipeline. **Fewer keypoints than you asked for.** `cv2.ORB_create(nfeatures=1000)` is a ceiling, not a promise; a low-texture image yields far fewer. **Grayscale is the convention.** Descriptors are computed on intensity, so colour information is discarded regardless. Convert explicitly with `cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)` so what you feed the detector is what you think it is — and remember `cv2.imread` gives you BGR, not RGB, in the first place. **The mask argument.** The second positional parameter is a detection mask (an 8-bit single-channel image, non-zero where detection is allowed), not an output. Passing `None` searches the whole image; passing a mask is the clean way to ignore a known-static region or a letterboxed border. **detect then compute can drop keypoints.** `compute()` may discard keypoints too close to the border to describe, so the list it returns can be shorter than the one you passed in — which is exactly why it returns the keypoint list too.
- Why does the row order of the descriptor array matter so much?Because matching is done on indices. A DMatch stores queryIdx and trainIdx, which are row numbers, and you convert them back to pixel coordinates with kp[m.queryIdx].pt. If you sort or filter one list without applying the identical operation to the other, every match still looks valid but points at the wrong location.
- What is the second argument to detectAndCompute, and when would you use it?It is a detection mask: an 8-bit single-channel image the same size as the input, non-zero where detection is permitted. Passing None searches everything. It is the clean way to skip a static overlay, a letterboxed border, or to restrict search to a tracked region of interest from the previous frame.
- Why is calling detect() then compute() sometimes different from detectAndCompute()?compute() can drop keypoints it cannot describe — typically those too near the image border for the descriptor patch to fit — which is why it returns a possibly shortened keypoint list alongside the descriptors. detectAndCompute also builds the scale-space pyramid once instead of twice, so it is faster.
saying these in an interview costs you the question
- Thinking detectAndCompute returns only descriptors
- Reading KeyPoint.pt as (row, column) like NumPy
- Assuming descriptors is an empty array when nothing is found
- Sorting keypoints without reordering the descriptor rows
- Treating nfeatures as a guaranteed keypoint count