skip to content

questions

3

A GraphQL response returns 200 with a populated errors array - what do you log?

level: seniorimportance: must knowfreq 58%

answer

  1. The status describes transport, not outcome
  2. Two error kinds, two different fields
  3. Path if it executed, locations if not
  4. Aggregate paths with indices stripped
  5. Internal cause on the line, not returned

basics

~20 s

Log every entry in the errors array with its path, or its document locations when it has no path, plus the error count, whether data came back partial or null, the document hash, and the internal cause.

solid answer

~50 s

The status is not the outcome, so the outcome has to be logged. For each entry in `errors`, record the machine-readable code from its `extensions` if the server sets one, and its position: a **field error** raised while resolving carries a `path` such as `enrollments.7.section.instructor`, while a **request error** from parsing or validation carries `locations` in the document and no path at all, because nothing executed. Record the count, and record whether `data` was complete, partial or `null` - a non-null field that errored propagates the null upward, so the shape of the hole is itself evidence. Then log the internal cause - the exception, the failing downstream call - which is exactly what the client's error message should not contain, tied to the response by a shared correlation id. Group by document hash to see whether one client or all of them are failing.

code

json · 10 lines
json
{
  "data": { "student": { "enrollments": [ { "section": null } ] } },
  "errors": [
    {
      "message": "Unable to resolve instructor",
      "path": ["student", "enrollments", 7, "section", "instructor"],
      "extensions": { "code": "UPSTREAM_UNAVAILABLE" }
    }
  ]
}

go deeper

for a junior

Remember that a GraphQL response can carry errors while the HTTP status still says 200, so success has to be read from the body. Know that an error entry may include a path naming the field that failed.

for a middle

Explain the difference between a request error from parsing or validation, which has locations and no path, and a field error raised during execution, which has a path and leaves data partially filled. Log the classification, not just a message.

for a senior

Demonstrate the investigation: aggregate index-stripped paths, group by document hash and client version, correlate the line to the client's response through a shared id, and know that failing validation is fast so no latency alert will fire.

for a principal

Own what 'failure' means for the platform: which classified errors count against a service objective when a partial response is a normal outcome, what every service must emit so incidents can be joined, and where the boundary sits between the client's message and the internal cause.

## The status code has stopped talking A GraphQL response is a map with `data`, `errors` and `extensions`. Under the long-standing `application/json` behaviour, a request that parsed and validated returns `200` whether every field resolved or none did; the GraphQL over HTTP specification - still a working draft - defines the newer `application/graphql-response+json` media type under which a request error may surface as a 4xx, but a *field* error still returns a success status with `errors` in the body. Either way the honest reading is: the status describes the transport, and the body describes the outcome. So the log line, not the status, has to carry the outcome. ## Two kinds of error, two different fields The specification distinguishes them, and confusing them wastes the first half of an incident. A **request error** happens before execution - the document did not parse, or failed validation against the schema, or the variables did not coerce. Nothing ran. `data` is absent entirely, and the error entry carries `locations` pointing into the document text. There is no `path`, because no field was reached. A **field error** happens during execution, when a resolver raises or returns something the schema forbids. Execution continues elsewhere, so `data` is present and partially filled, and the entry carries a `path`: the list of response keys and list indices leading to the failing field, for example `["student", "enrollments", 7, "section", "instructor"]`. A log line that flattens both into `errorMessage` throws away the single most useful discriminator you have. Record a classification field, the `path` when present, the `locations` when not. ## Why the paths are the payoff Paths aggregate. In a course-enrolment graph, a 4-person platform team seeing 1,184 failures in ten minutes learns nothing from 1,184 copies of "Cannot fetch instructor". Grouped by path with the list indices stripped, they learn that every one of them is `student.enrollments.*.section.instructor` and none are anywhere else - which points at one resolver and one downstream, immediately. Paths also expose null propagation. If `Section.instructor` is declared non-null and it errors, execution cannot leave a null there, so the null propagates to the nearest nullable ancestor; the client may receive a `null` section, or a `null` enrollment, while the error path still names the field that actually failed. Logging both the error path and whether `data` was partial or fully `null` is what lets you explain to the mobile team why a whole timetable vanished over one missing instructor. ## The internal detail belongs here, not there The response goes to a caller you do not control, so its messages should be safe and generic. The log line goes to your team, so it is the correct home for the exception type, the failing downstream call, the identifier that was not found. The two are joined by a **correlation id** carried on the line and, by convention, echoed to the client in the response `extensions`, so a screenshot in a support ticket becomes a log query. ## Working the failure Take the shape this leaf is named for. Mobile build 7.3.1 ships on a Tuesday; a schema change renamed a field the build still selects. From that afternoon, a slice of traffic returns a response with no `data` at all, one error, `locations` at line 4 column 7, no path. Because the line carries the document hash, the picture resolves in one query: every failure shares a single hash, the hash appears only from clients reporting version 7.3.1, and the rate is flat rather than growing - a fixed population of installs, not a spreading fault. No latency alert fires, because failing validation is fast, and no status-code alert fires either under the legacy behaviour. The alert that catches it is one built on the logged error count. That is the case for logging errors as a *counted, classified, path-keyed* field rather than as a message: it makes the difference between "one broken client build" and "the enrolment service is down" visible in seconds. ## Fields worth having, concretely Errors count. A classification per entry. The path, index-normalized, when there is one. Locations when there is not. The `extensions` code, if the server assigns one. Whether `data` was complete, partial or null. The document hash. The correlation id. The internal cause. And, deliberately, not the raw variables and not the document text. One caveat: an error count on a line is not by itself an error rate. Some schemas use nullable fields and error entries as a normal, expected outcome, so a blanket alert on "any error" is noise. Alert on classified errors per document hash, and let the shape of the baseline tell you what normal looks like.

  • Why can an error's path point at a field while the client sees a null much higher up?
    Because a non-null field cannot hold a null. When a field declared non-null errors, execution cannot put a null in that position, so the null is propagated to the nearest nullable ancestor - potentially several levels up, or to the whole `data` entry if every ancestor is non-null. The error entry still names the field that actually failed, which is why logging the path alongside how much of `data` survived is what lets you explain a disproportionately large hole.
  • How do you turn thousands of error paths into something you can alert on?
    Normalize before grouping. Strip the list indices, so `enrollments.7.section.instructor` and `enrollments.0.section.instructor` collapse to one key, and group that key together with the document hash and the error classification. What you then alert on is a classified error count per group crossing a threshold, not the presence of any error at all - some schemas return error entries as an ordinary outcome, and a blanket alert on those is permanent noise.
  • A client reports a failure but you have no request id from them. What in the line still finds it?
    The document hash narrows it to one operation shape, the client name and version narrow it to their build, and the authenticated principal id narrows it to their session; a timestamp window closes the gap. That combination is usually enough. It is also the argument for echoing a correlation id into the response `extensions` by convention, so the next report arrives with the identifier already attached rather than being reconstructed.
  • Should the log line's error message be the same string the client received?
    No - they serve different readers. The client's message should be safe and generic because it reaches a caller you do not control; the line should carry the real cause, the exception type and the failing downstream call. Log both if you like, but the point of the split is that the detailed one exists only on your side, joined to the client's copy by a shared correlation id.

The 200 is the courier confirming the parcel arrived. The errors array is the packing note saying two of the eleven items are missing, and the error path is the line number on that note telling you which two.

saying these in an interview costs you the question

  • Treating a 200 status as a successful outcome
  • Logging only the error message, dropping the path
  • Confusing validation errors with resolver failures
  • Expecting a path on a parse or validation error
  • Alerting on any error entry without classification
  • Putting the internal cause into the client response

context

open as a page

What does a per-operation GraphQL log line record, and what stays off it?

level: juniorimportance: should knowfreq 44%

basics

~20 s

A per-operation GraphQL log line records the operation type and name, a hash of the executed document, the duration, and the path of every error the response carried. Raw variable values and the full document text stay off it.

open as a page

Why is a GraphQL operation name an unreliable key for grouping log lines?

level: middleimportance: nice to knowfreq 28%

basics

~10 s

The operation name is arbitrary text the client writes in its own document. Nothing binds it to that document, different documents can share one name, and an operation may have no name at all.

open as a page