How do you turn Gemini's groundingMetadata into inline citations in your UI?
answer
- Two lists joined by indices
- Segment offsets are byte offsets
- Chunk indices point into groundingChunks
- Redirect URIs expire, don't archive them
- Search Suggestions HTML must be shown
basics
~20 sWalk groundingSupports: each has a text segment with start and end indices and a list of chunk indices. Slice the answer by those offsets to place markers, resolve the indices against groundingChunks for the source titles and URIs, and render the searchEntryPoint HTML as required.
solid answer
~50 sEach entry in `groundingSupports` gives you a `segment` — start index, end index and the quoted text — plus `groundingChunkIndices`, which index into the `groundingChunks` list where each web chunk carries a `uri` and `title`. Rendering is: sort the supports by start offset, insert a footnote marker at each segment's end, and build the footnote from the chunks the indices point at. Two traps. First, the segment offsets are **byte** offsets into the part's text, so slicing a language-level string by them mis-aligns as soon as the answer contains non-ASCII characters — encode to UTF-8, slice, decode back. Second, the chunk URIs are Google redirect links documented as valid only for a limited period, so they are display links, not archival citations: resolve or snapshot what you need to keep. Finally, `searchEntryPoint.renderedContent` is HTML you are required to display alongside grounded answers, and when streaming, the metadata generally arrives with the closing chunks, so render citations after the stream settles.
code
python · 25 linesfrom google import genai
from google.genai import types
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="Summarise this week's changes to the tax filing deadline.",
config=types.GenerateContentConfig(
tools=[types.Tool(google_search=types.GoogleSearch())],
),
)
candidate = response.candidates[0]
metadata = candidate.grounding_metadata
if metadata and metadata.grounding_supports:
raw = candidate.content.parts[0].text.encode("utf-8")
for support in metadata.grounding_supports:
segment = support.segment
quoted = raw[segment.start_index : segment.end_index].decode("utf-8")
sources = [
metadata.grounding_chunks[i].web.uri
for i in support.grounding_chunk_indices or []
]
print(quoted, "<-", sources)go deeper
Know that citations come from groundingSupports and groundingChunks rather than from parsing the answer text, and that each support names both a span and the sources behind it.
Explain the join: supports carry a segment plus chunk indices into the chunk list, offsets are byte-based, and some spans legitimately have no support at all.
Demonstrate the production details — byte-space slicing, stable footnote numbering, the temporary redirect URIs, the mandatory Search Suggestions rendering, and applying markers only after a stream settles.
Set the provenance policy: what the product claims a citation means, whether source URLs are resolved and retained for audit, and how attribution is presented so users are not misled about which claims are actually backed.
## The data model Grounding attribution in Gemini is expressed as two lists plus a join. `groundingChunks` is the flat source list. For web grounding each element has a `web` object with `uri` and `title`. Positions in this list are meaningful — they are the keys the other list refers to. `groundingSupports` is the join. Each entry has a `segment` describing a contiguous piece of the generated answer (with `startIndex`, `endIndex` and the `text` itself) and a `groundingChunkIndices` array of integers pointing into `groundingChunks`. One sentence can be supported by several chunks; one chunk can support several sentences; and — importantly — plenty of the answer may have no support entry at all, because the model wrote it from its own knowledge rather than from a retrieved page. That last fact shapes the UI. If you render "every sentence has a footnote", you are lying about the unattributed parts. The honest rendering marks only the spans that actually carry supports. ## The byte-offset trap `startIndex` and `endIndex` are measured in **bytes** over the UTF-8 encoding of the part's text, not in characters or code units. In an ASCII-only English answer nobody notices. The first time the answer contains an accented name, a currency symbol, an em dash or any non-Latin script, naive string slicing drifts and your footnote markers land mid-word — a bug that passes every test written in English and fails immediately in production. The fix is mechanical: encode the text to UTF-8 once, slice by the given indices in byte space, decode the slice back for display, and compute your marker insertion points in the same byte space. Do the whole citation pass in bytes and convert only at the end. ## Rendering the markers A workable algorithm: 1. Take the candidate's text and encode it to UTF-8. 2. Collect the supports and sort them by `startIndex`. 3. Assign each distinct chunk index a stable footnote number in order of first appearance, so the numbering reads sequentially down the answer rather than reflecting the arbitrary order of the chunk list. 4. Walk the supports from the *end* backwards and insert markers at each `endIndex`, so earlier insertions do not shift the offsets of later ones — or build the output by copying spans forward. Either works; mixing them is where off-by-one bugs live. 5. Decode and render, then emit the footnote list with each chunk's `title` linked to its `uri`. Overlapping supports happen; decide a policy (merge markers, or keep the first) rather than letting them stack into unreadable clutter. ## The URIs are temporary The `uri` on a web chunk is not the publisher's canonical URL — it is a Google grounding redirect, and the documentation describes it as accessible only for a limited window (on the order of weeks) after the response. The consequences are concrete: - Do not store these links as permanent citations in an audit log or an exported report; they will rot. - Do not present them as the source's own domain — use the returned `title` for the label so the user sees what they are clicking. - If you need durable provenance, resolve the redirect at response time and store the final URL, or snapshot the retrieved content, subject to whatever your legal position on caching is. ## The display requirement `searchEntryPoint.renderedContent` contains ready-made HTML and CSS for the Google Search Suggestions chips. Displaying it with grounded answers is a condition of using the feature, not an optional nicety, and it is a compliance item your reviewers will ask about. Practically, that means your renderer needs a slot for injected HTML — which in a React or similar app means a deliberate, scoped dangerous-HTML insertion for this one trusted field, kept away from any path that renders user-supplied markup. ## Streaming With `streamGenerateContent` the text arrives in chunks, and grounding metadata generally shows up with the closing chunks rather than the first. So the natural UI is: stream the prose immediately, then apply citation markers once the stream completes and the metadata is in hand. Attempting to attach markers to partial text is fragile — the offsets refer to the full generated part, and a marker computed against a half-received string will land in the wrong place. Accumulate the metadata from whichever chunk carries it and do one citation pass at the end. ## When there is nothing to cite If the model did not search, `groundingMetadata` is absent entirely, and if it searched but produced no supports you may get chunks without a span mapping. Both are normal. Your renderer needs a graceful path: show the answer plainly, show the sources as a plain list if you have chunks but no supports, and never crash on a missing field. The most common production bug in this whole area is a null-pointer on `grounding_metadata` for a request the model simply answered from memory.
- Why does slicing the answer by segment indices break on non-English text?Because the indices are byte offsets into the UTF-8 encoding, while most languages slice strings by characters or UTF-16 code units. Any multi-byte character — an accent, an em dash, CJK text — shifts every subsequent character position, so markers land mid-word and quoted spans come back garbled. Encode to bytes, slice there, decode the result, and keep all offset arithmetic in byte space.
- Can you store the chunk URIs as permanent citations in an audit log?No. They are Google grounding redirect links documented as valid only for a limited period, so an audit record built from them decays into dead links. If you need durable provenance, resolve the redirect when the response arrives and store the destination URL, or snapshot the content within whatever your legal position on caching allows, and keep the returned title as the human-readable label.
- Parts of a grounded answer have no support entries at all. What should the UI do?Leave them unmarked. Absence of a support means the model wrote that span from its own knowledge rather than a retrieved page, and marking it anyway misrepresents the attribution. Render markers only on supported spans; if you have chunks but no supports at all, show the sources as a plain consulted-list without implying sentence-level backing.
- When streaming a grounded response, when do you apply the citation markers?After the stream finishes. Metadata generally arrives with the closing chunks, and the segment offsets refer to the complete generated part, so computing marker positions against partially received text puts them in the wrong place. Stream the prose for responsiveness, accumulate the metadata from whichever chunk carries it, then run one citation pass over the settled text.
saying these in an interview costs you the question
- Slices the answer by character index instead of byte offsets
- Stores grounding redirect URIs as permanent citations
- Assumes every sentence has a grounding support entry
- Skips rendering the searchEntryPoint Search Suggestions HTML
- Tries to attach citation markers to partial streamed text