In Cohere's v2/chat API, what do the p and k parameters control?
answer
- Short names, not the OpenAI ones
- Nucleus and rank cutoffs
- Zero means off for one of them
- Default is cooler than you expect
- Ceiling of one on this endpoint
basics
~20 sCohere spells nucleus and top-k sampling as p and k, not top_p and top_k. Both narrow the candidate pool the next token is drawn from; k defaults to 0, which disables top-k. Temperature is separate and defaults to 0.3.
solid answer
~40 sOn `POST /v2/chat`, Cohere's randomness controls are `temperature`, `p` and `k` — the short names are the vendor detail that trips people up. `p` is nucleus sampling: the model samples only from the smallest set of tokens whose probabilities sum to `p`. `k` is top-k: only the `k` most likely tokens are considered, and it defaults to `0`, meaning top-k filtering is off. `temperature` defaults to `0.3` on this endpoint, which is lower than most vendors' defaults, so a port from another provider can look unexpectedly conservative. There are also `seed`, `stop_sequences`, `frequency_penalty` and `presence_penalty`. The practical trap is pasting an OpenAI-shaped body: `top_p` and `top_k` are not the names Cohere reads, so your intended setting does not take effect.
code
python · 14 linesimport cohere
co = cohere.ClientV2(api_key="COHERE_API_KEY")
res = co.chat(
model="command-a-03-2025",
messages=[{"role": "user", "content": "List three uses for a bloom filter."}],
temperature=0.2,
p=0.9,
k=0,
stop_sequences=["###"],
)
print(res.message.content[0].text)go deeper
Recall the field names: temperature, p and k on Cohere's chat endpoint, with p and k narrowing which tokens can be picked. Say that k of 0 turns top-k off.
Explain how the knobs stack — temperature reshapes, p and k truncate — and name the defaults, including the 0.3 temperature default and the 1.0 cap, plus what happens to a top_p field sent to Cohere.
Demonstrate the migration discipline: one tuned dial, an explicit per-vendor parameter adapter, and tests that catch untranslated field names before a silent quality regression reaches users.
Own the policy. Decide which sampling settings are product-level defaults versus per-feature overrides, how safety_mode changes are governed and re-evaluated, and how the platform keeps generation settings comparable across multiple model vendors.
## The names are the exam question The decoding *theory* here is vendor-neutral: temperature reshapes the probability distribution, nucleus sampling truncates it by cumulative mass, top-k truncates it by rank. What is specific to Cohere — and what an interviewer is probing — is that the Command chat endpoint spells these as **`temperature`, `p` and `k`**, not `top_p` and `top_k`. Cohere carried the short names over from its original generate API into `/v2/chat`, so a request body copied from an OpenAI example and pointed at Cohere sets fields Cohere is not reading. ## The fields on /v2/chat - **`temperature`** — defaults to `0.3` and is capped at `1.0` on the chat endpoint. Two things matter here. First, the default is noticeably lower than the ~1.0 default several other vendors ship, so the same prompt will feel more deterministic on Cohere out of the box. Second, the ceiling is 1.0, so a config that allows 1.5 or 2.0 elsewhere has no equivalent here and will be rejected or clamped rather than producing wilder output. - **`p`** — nucleus sampling. Its documented default sits below 1.0, so mild nucleus truncation is active even if you never set it. Lower values tighten the pool. - **`k`** — top-k. Defaults to `0`, which means the filter is **disabled**, not that it selects zero tokens. This asymmetry with `p` is a favourite quiz item: `k = 0` is "off", while `p = 0` would be meaningless. - **`seed`** — a best-effort determinism hint. Same seed plus identical request tends to produce the same output, but it is not a contractual guarantee across model snapshots or infrastructure changes; never build a test suite that asserts exact strings on it. - **`stop_sequences`** — strings that end generation; a hit is reported back as `finish_reason: "STOP_SEQUENCE"`. - **`frequency_penalty` / `presence_penalty`** — repetition controls. On Cohere these live in a 0.0–1.0 band rather than the −2.0–2.0 band some other APIs use, another silent porting mismatch. - **`safety_mode`** — a Cohere-specific enum (`CONTEXTUAL`, `STRICT`, `OFF`, with availability depending on the model) that adjusts the built-in safety preamble. It has no analogue in most OpenAI-shaped payloads. ## Combining temperature with p and k The controls stack: temperature reshapes the distribution, then `p` and `k` truncate it, then a token is drawn. Turning all three knobs at once makes behaviour hard to reason about, and it makes incident forensics harder — when output quality regresses you want one variable to have moved. The usual production discipline is to fix `p` and `k` at their defaults and treat `temperature` as the single tuned dial, or, for extraction and classification work, drive temperature to `0` and stop tuning altogether. It is also worth remembering that low temperature does not equal reproducibility. Batching, floating-point non-determinism on GPU kernels and snapshot changes all mean two identical requests can differ. If your product needs stable outputs, cache the response, do not re-derive it. ## Porting from an OpenAI-shaped payload This is the concrete failure mode. A team migrates a service, keeps the request builder, and swaps the base URL and key. `top_p: 0.2` goes over the wire; Cohere either rejects the unknown field or ignores it, and `p` stays at its default. The output distribution silently widens relative to what was tested, and nobody notices until a quality report drifts. The mitigations are ordinary engineering: map parameters explicitly in a per-vendor adapter rather than passing a shared dict through, assert on the echoed request where the API supports it, and add a smoke test that fails loudly when a field the adapter is supposed to translate reaches the wire untranslated. Cohere does publish an OpenAI-compatibility layer that accepts the OpenAI request shape and translates it, which is a legitimate migration path — but it is a different surface with its own coverage gaps, so choose it deliberately instead of assuming the native endpoint is lenient. ## What to say in an interview Name the three fields, say `p` and `k` rather than `top_p`/`top_k`, mention that `k = 0` disables top-k and that temperature defaults low at 0.3 and is capped at 1.0, then pivot to the operational point: pick one dial, keep the rest at defaults, and never assume parameter names or ranges survive a provider swap. That last sentence is what separates an answer that recites documentation from one that sounds like it has run a migration.
- What does safety_mode do on a Cohere chat request, and when would you change it?`safety_mode` selects which built-in safety instruction Cohere prepends. `CONTEXTUAL` is the general-purpose setting, `STRICT` tightens refusals for sensitive deployments, and `OFF` is available only on some models and only where your agreement allows it. Change it deliberately — it alters refusal behaviour, so any change needs a re-run of your safety evals, not just a spot check.
- Your team wants reproducible outputs for a regression suite. Is seed enough?No. `seed` is best effort: it makes repeats likely, not guaranteed, and it does not survive a model snapshot change, so a suite asserting exact strings will break on the next upgrade. Pin the dated model ID, set temperature to 0 to reduce variance, and assert on properties — schema validity, required fields, semantic similarity — rather than byte equality. Cache golden outputs if you truly need fixed text.
- How do you keep parameter mismatches from reaching production when you support several vendors?Put a per-vendor adapter between your internal generation settings and the wire body, so a `nucleus` setting maps to `p` for Cohere and `top_p` elsewhere and nothing is passed through blindly. Then add a contract test per vendor that serialises a known settings object and asserts the exact field names and ranges emitted. Mismatches then fail in CI rather than as a slow quality drift.
saying these in an interview costs you the question
- Calling them top_p and top_k on Cohere's endpoint
- Reading k=0 as blocking all tokens rather than disabling top-k
- Assuming temperature defaults to 1.0 like other vendors
- Setting temperature above 1.0 on the chat endpoint
- Treating seed as a hard determinism guarantee