skip to content

In Cohere's v2/chat API, what do the p and k parameters control?

level: middleimportance: should knowfreq 45%

answer

  1. Short names, not the OpenAI ones
  2. Nucleus and rank cutoffs
  3. Zero means off for one of them
  4. Default is cooler than you expect
  5. Ceiling of one on this endpoint

basics

~20 s

Cohere spells nucleus and top-k sampling as p and k, not top_p and top_k. Both narrow the candidate pool the next token is drawn from; k defaults to 0, which disables top-k. Temperature is separate and defaults to 0.3.

solid answer

~40 s

On `POST /v2/chat`, Cohere's randomness controls are `temperature`, `p` and `k` — the short names are the vendor detail that trips people up. `p` is nucleus sampling: the model samples only from the smallest set of tokens whose probabilities sum to `p`. `k` is top-k: only the `k` most likely tokens are considered, and it defaults to `0`, meaning top-k filtering is off. `temperature` defaults to `0.3` on this endpoint, which is lower than most vendors' defaults, so a port from another provider can look unexpectedly conservative. There are also `seed`, `stop_sequences`, `frequency_penalty` and `presence_penalty`. The practical trap is pasting an OpenAI-shaped body: `top_p` and `top_k` are not the names Cohere reads, so your intended setting does not take effect.

code

python · 14 lines
python
import cohere

co = cohere.ClientV2(api_key="COHERE_API_KEY")

res = co.chat(
    model="command-a-03-2025",
    messages=[{"role": "user", "content": "List three uses for a bloom filter."}],
    temperature=0.2,
    p=0.9,
    k=0,
    stop_sequences=["###"],
)

print(res.message.content[0].text)

go deeper

for a junior

Recall the field names: temperature, p and k on Cohere's chat endpoint, with p and k narrowing which tokens can be picked. Say that k of 0 turns top-k off.

for a middle

Explain how the knobs stack — temperature reshapes, p and k truncate — and name the defaults, including the 0.3 temperature default and the 1.0 cap, plus what happens to a top_p field sent to Cohere.

for a senior

Demonstrate the migration discipline: one tuned dial, an explicit per-vendor parameter adapter, and tests that catch untranslated field names before a silent quality regression reaches users.

for a principal

Own the policy. Decide which sampling settings are product-level defaults versus per-feature overrides, how safety_mode changes are governed and re-evaluated, and how the platform keeps generation settings comparable across multiple model vendors.

## The names are the exam question The decoding *theory* here is vendor-neutral: temperature reshapes the probability distribution, nucleus sampling truncates it by cumulative mass, top-k truncates it by rank. What is specific to Cohere — and what an interviewer is probing — is that the Command chat endpoint spells these as **`temperature`, `p` and `k`**, not `top_p` and `top_k`. Cohere carried the short names over from its original generate API into `/v2/chat`, so a request body copied from an OpenAI example and pointed at Cohere sets fields Cohere is not reading. ## The fields on /v2/chat - **`temperature`** — defaults to `0.3` and is capped at `1.0` on the chat endpoint. Two things matter here. First, the default is noticeably lower than the ~1.0 default several other vendors ship, so the same prompt will feel more deterministic on Cohere out of the box. Second, the ceiling is 1.0, so a config that allows 1.5 or 2.0 elsewhere has no equivalent here and will be rejected or clamped rather than producing wilder output. - **`p`** — nucleus sampling. Its documented default sits below 1.0, so mild nucleus truncation is active even if you never set it. Lower values tighten the pool. - **`k`** — top-k. Defaults to `0`, which means the filter is **disabled**, not that it selects zero tokens. This asymmetry with `p` is a favourite quiz item: `k = 0` is "off", while `p = 0` would be meaningless. - **`seed`** — a best-effort determinism hint. Same seed plus identical request tends to produce the same output, but it is not a contractual guarantee across model snapshots or infrastructure changes; never build a test suite that asserts exact strings on it. - **`stop_sequences`** — strings that end generation; a hit is reported back as `finish_reason: "STOP_SEQUENCE"`. - **`frequency_penalty` / `presence_penalty`** — repetition controls. On Cohere these live in a 0.0–1.0 band rather than the −2.0–2.0 band some other APIs use, another silent porting mismatch. - **`safety_mode`** — a Cohere-specific enum (`CONTEXTUAL`, `STRICT`, `OFF`, with availability depending on the model) that adjusts the built-in safety preamble. It has no analogue in most OpenAI-shaped payloads. ## Combining temperature with p and k The controls stack: temperature reshapes the distribution, then `p` and `k` truncate it, then a token is drawn. Turning all three knobs at once makes behaviour hard to reason about, and it makes incident forensics harder — when output quality regresses you want one variable to have moved. The usual production discipline is to fix `p` and `k` at their defaults and treat `temperature` as the single tuned dial, or, for extraction and classification work, drive temperature to `0` and stop tuning altogether. It is also worth remembering that low temperature does not equal reproducibility. Batching, floating-point non-determinism on GPU kernels and snapshot changes all mean two identical requests can differ. If your product needs stable outputs, cache the response, do not re-derive it. ## Porting from an OpenAI-shaped payload This is the concrete failure mode. A team migrates a service, keeps the request builder, and swaps the base URL and key. `top_p: 0.2` goes over the wire; Cohere either rejects the unknown field or ignores it, and `p` stays at its default. The output distribution silently widens relative to what was tested, and nobody notices until a quality report drifts. The mitigations are ordinary engineering: map parameters explicitly in a per-vendor adapter rather than passing a shared dict through, assert on the echoed request where the API supports it, and add a smoke test that fails loudly when a field the adapter is supposed to translate reaches the wire untranslated. Cohere does publish an OpenAI-compatibility layer that accepts the OpenAI request shape and translates it, which is a legitimate migration path — but it is a different surface with its own coverage gaps, so choose it deliberately instead of assuming the native endpoint is lenient. ## What to say in an interview Name the three fields, say `p` and `k` rather than `top_p`/`top_k`, mention that `k = 0` disables top-k and that temperature defaults low at 0.3 and is capped at 1.0, then pivot to the operational point: pick one dial, keep the rest at defaults, and never assume parameter names or ranges survive a provider swap. That last sentence is what separates an answer that recites documentation from one that sounds like it has run a migration.

  • What does safety_mode do on a Cohere chat request, and when would you change it?
    `safety_mode` selects which built-in safety instruction Cohere prepends. `CONTEXTUAL` is the general-purpose setting, `STRICT` tightens refusals for sensitive deployments, and `OFF` is available only on some models and only where your agreement allows it. Change it deliberately — it alters refusal behaviour, so any change needs a re-run of your safety evals, not just a spot check.
  • Your team wants reproducible outputs for a regression suite. Is seed enough?
    No. `seed` is best effort: it makes repeats likely, not guaranteed, and it does not survive a model snapshot change, so a suite asserting exact strings will break on the next upgrade. Pin the dated model ID, set temperature to 0 to reduce variance, and assert on properties — schema validity, required fields, semantic similarity — rather than byte equality. Cache golden outputs if you truly need fixed text.
  • How do you keep parameter mismatches from reaching production when you support several vendors?
    Put a per-vendor adapter between your internal generation settings and the wire body, so a `nucleus` setting maps to `p` for Cohere and `top_p` elsewhere and nothing is passed through blindly. Then add a contract test per vendor that serialises a known settings object and asserts the exact field names and ranges emitted. Mismatches then fail in CI rather than as a slow quality drift.

saying these in an interview costs you the question

  • Calling them top_p and top_k on Cohere's endpoint
  • Reading k=0 as blocking all tokens rather than disabling top-k
  • Assuming temperature defaults to 1.0 like other vendors
  • Setting temperature above 1.0 on the chat endpoint
  • Treating seed as a hard determinism guarantee

context