LLM Providers & Models
The model vendors you actually call from code — OpenAI, Anthropic, Gemini, Mistral, Cohere, DeepSeek, Qwen — plus the gateways that sit in front of them. Interviews here probe whether you know each vendor's API shape, model line-up, and cost/limit behaviour; the model theory underneath is taught once in the foundations trees.
on this pageshowhide
explore
- OpenAI35 questions
- Chat Completions API6 questions
- Tool & Function Calling6 questions
- Structured Outputs5 questions
- Embeddings API6 questions
- Assistants & Responses API6 questions
- Rate Limits & Cost6 questions
- Anthropic31 questions
- Messages API fundamentals5 questions
- Claude model tiers5 questions
- Tool use and function calling5 questions
- Streaming responses via SSE6 questions
- Prompt caching5 questions
- Vision and multimodal inputs5 questions
- Google Gemini32 questions
- generateContent API5 questions
- Multimodal Inputs5 questions
- Function Calling6 questions
- Long-Context & Caching5 questions
- Embeddings5 questions
- Safety & Grounding6 questions
- Cohere24 questions
- Command model family6 questions
- Grounded generation and citations6 questions
- Embed API6 questions
- Rerank API6 questions
- Mistral19 questions
- Model family and licensing5 questions
- Chat completions API5 questions
- Function calling4 questions
- Embeddings and specialised endpoints5 questions
- xAI5 questions
- DeepSeek19 questions
- V3 and R1 model family5 questions
- OpenAI-compatible API5 questions
- Reasoning model API behaviour5 questions
- Pricing and context caching4 questions
- Meta Llama34 questions
- Model Family and Sizes5 questions
- Prompt and Chat Templates5 questions
- Quantization6 questions
- Fine-Tuning6 questions
- Self-Hosted Inference6 questions
- Licensing and Compliance6 questions
- Google Gemma5 questions
- Qwen15 questions
- Model family and variants5 questions
- Hosted API (Model Studio / DashScope)5 questions
- Self-hosting Qwen weights5 questions
- Opus3 questions
- Sonnet3 questions
- OpenAI Image Generation6 questions
- Whisper6 questions
- OpenRouter17 questions
- Unified OpenAI-compatible API5 questions
- Routing, fallbacks and provider selection6 questions
- Cost, credits and usage accounting6 questions
- GitHub Models6 questions
questions
260 · 16 sectionsIn OpenAI Chat Completions, what do the system, user, and assistant roles do?
basics
~20 sEvery entry in the messages array carries a role. system holds standing instructions the model should follow throughout, user holds what the human said, and assistant holds the model's own earlier replies that you replay as context.
What does OpenAI's /v1/embeddings endpoint take as input and return?
basics
~20 sA POST to /v1/embeddings sends a model plus input (a string or an array of strings) and returns a data array of vectors, each carrying its embedding and the index of the input it came from, plus a usage object counting input tokens.
How is an OpenAI API call billed, and why do long chats cost more?
basics
~20 sOpenAI bills per token, quoted per million tokens, and output tokens cost several times more than input tokens. The API is stateless, so every turn resends the whole conversation as input — the input bill grows with each turn.
In OpenAI's Responses API, what does previous_response_id actually do?
basics
~20 sprevious_response_id points at a stored earlier response, so the server replays that whole conversation as the prefix of the new request. You send only the new turn, but you are still billed for every replayed input token.
How do you make documents searchable by OpenAI's hosted file_search tool?
basics
~20 sUpload the files, add them to a vector store, and wait for indexing to finish, then reference that store's id from the file_search tool on your request. OpenAI handles chunking, embedding, and ranking; you pay for storage per day.
In Anthropic's Messages API, how do you send a system prompt to Claude?
basics
~20 sThe Messages API takes the system prompt as a top-level system parameter on the request body, alongside model, max_tokens and messages. It is not a message inside the messages array; that array carries only user and assistant turns.
What do the Haiku, Sonnet, and Opus tiers mean in Claude's model line-up?
basics
~20 sHaiku, Sonnet and Opus are Anthropic's three Claude size tiers, ordered smallest to largest. Haiku is the fastest and cheapest, Opus is the most capable and most expensive, and Sonnet sits between them. All three are called through the same Messages API.
How do you declare a tool for Claude in Anthropic's Messages API?
basics
~10 sSend a top-level tools array on the Messages request. Each entry needs a name, a description that tells Claude when to call it, and an input_schema: a JSON Schema object describing the parameters.
How do you send an image to Claude's Messages API inside a user message?
basics
~20 sSet the user message content to an array of blocks and add an image block next to your text block. The image block carries a source object that is either base64 (with media_type and data) or a url. Image blocks belong in user turns.
What does the required max_tokens parameter cap on an Anthropic Messages API call?
basics
~20 smax_tokens is a hard ceiling on the tokens Claude generates in that one response, including any thinking tokens. It is required on every Messages API request, and hitting it truncates the output mid-stream with stop_reason "max_tokens".
In Google's Gemini API, what does the model return when it decides to call a declared function?
basics
~20 sGemini returns an ordinary response whose candidate content holds a FunctionCall part: a function name plus an args object of already-parsed arguments. The model never runs your code — your application executes the call and sends the result back.
How do you make a basic Gemini generateContent call with the google-genai SDK?
basics
~10 sCreate a client with genai.Client(), which picks up the GEMINI_API_KEY environment variable, then call client.models.generate_content(model=..., contents=...). You get back a GenerateContentResponse whose .text property joins the text parts of the first candidate.
How do you send an image to Gemini's generateContent as inline data or a file URI?
basics
~20 sGemini takes media as a Part inside contents. Small files ride along as inline_data (raw bytes plus a mime_type); larger or reused files are uploaded to the Files API first and referenced as file_data with a file_uri and mime_type.
In the Gemini API, what does a safetySettings entry with a category and threshold control?
basics
~20 sEach safetySettings entry pairs a harm category (harassment, hate speech, sexually explicit, dangerous content, civic integrity) with a block threshold. The threshold sets how confident the filter must be before Gemini refuses to return content in that category.
What does task_type change in a Gemini embedContent request?
basics
~20 stask_type tells Gemini's embedding model what the text is for — a stored passage, a search query, a classification input — and the model places the same text differently in vector space for each. Index and query sides must use the matching pair.
What does a Cohere v2/chat request require, and what shape is the response?
basics
~20 sA Cohere v2/chat call needs a model ID and a messages array of role/content objects, authenticated with a bearer API key. The response carries an assistant message whose content is a list of typed blocks, plus finish_reason and usage.
How do you call Cohere's v2/embed endpoint and read the vectors from the response?
basics
~20 sSend texts plus a model and an input_type to POST /v2/embed. The response groups vectors by type: a float request returns them under embeddings.float, one vector per input text, in the order you sent them.
In Cohere's v2 Chat API, how do you pass source documents so the reply carries citations?
basics
~20 sSend a top-level documents array alongside messages in the v2 chat call. Each entry is a plain string or an object with an id and a data map. The reply's message.citations then link answer spans to those documents.
What does a Cohere v2/rerank request contain, and what does the API return?
basics
~20 sA Cohere v2/rerank call sends a model name, one query string, and a documents list of plain strings, plus an optional top_n. It returns a results array ordered by relevance, each entry carrying the document's original index and a relevance_score.
Which models make up Cohere's Command family, and why pin a dated model ID?
basics
~20 sCommand A is Cohere's current flagship chat model with a 256K-token context, above the earlier Command R+ and Command R at 128K and the small Command R7B. IDs carry a date suffix; undated aliases follow the newest snapshot, so production should pin the dated ID.
What fields does a minimal Mistral chat completions request require?
basics
~10 sMistral's POST /v1/chat/completions needs only two body fields: model (for example mistral-large-latest) and messages, an ordered array of role/content objects. Authentication is a bearer API key; max_tokens, temperature and the rest are optional.
How do you call Mistral's /v1/embeddings endpoint, and what does mistral-embed return?
basics
~20 sPOST /v1/embeddings on api.mistral.ai with a bearer API key, model "mistral-embed" and an "input" array of strings. The response carries a data array of 1024-dimension float vectors, each tagged with the index of the string it came from, plus a token usage block.
In Mistral's chat completions response, which fields carry a tool call?
basics
~10 sMistral sets choices[0].finish_reason to "tool_calls" and fills choices[0].message.tool_calls. Each entry has an id, a function.name, and function.arguments — a JSON-encoded string you must parse before running your code.
How do you assemble a full reply from Mistral's streamed chat completion chunks?
basics
~20 sSet stream true on Mistral's chat completions request and the reply arrives as chat.completion.chunk objects over server-sent events. Concatenate choices[0].delta.content across chunks in arrival order; the last chunk carries finish_reason, and the stream closes with data: [DONE].
Porting tool_choice "required" from OpenAI to Mistral — which value do you use?
basics
~20 sUse "any". Mistral's documented tool_choice values are auto (the default when tools are present), any (the model must call a tool), and none (tools stay declared but unused); "required" is OpenAI's spelling for the any mode.
How do you call xAI's Grok models using the OpenAI SDK instead of OpenAI's?
basics
~20 sxAI exposes an OpenAI-compatible chat completions API. Keep the OpenAI SDK, point base_url at https://api.x.ai/v1, authenticate with your xAI key (conventionally XAI_API_KEY) as the bearer token, and pass a Grok model id such as grok-4.
How do you enable xAI's Live Search in a Grok chat completions request?
basics
~20 sSend a search_parameters object on the chat completions request. Its mode field turns search on, off, or leaves the decision to the model; sources selects web, x, news or rss. The response carries a citations array, and each source used is billed on top of tokens.
How does xAI's API expose reasoning effort and reasoning traces on Grok?
basics
~20 sOn xAI's lighter Grok reasoning models you set reasoning_effort to "low" or "high" to trade depth against latency and cost. The trace comes back in the message as reasoning_content, and thinking tokens are billed as completion tokens with a reasoning_tokens breakdown in usage.
Which OpenAI chat-completion parameters does xAI's Grok 4 family reject?
basics
~20 sxAI documents presence_penalty, frequency_penalty, stop and reasoning_effort as unsupported on the grok-4 family: sending them fails the request rather than being ignored. Compatibility is per model, so a shared parameter struct must be filtered before dispatch.
For a production Grok service, when is xAI's native SDK worth it over OpenAI-compatible calls?
basics
~20 sUse the OpenAI-compatible endpoint when portability across vendors is the priority and you mostly need plain chat. Reach for xai-sdk when xAI-specific capabilities — live search, deferred sampling, typed helpers — are central enough that untyped passthrough fields become a liability.
How do you call DeepSeek's API using the OpenAI SDK, and what must you change?
basics
~10 sPoint the OpenAI client at base_url https://api.deepseek.com, pass your DeepSeek key as api_key, and use a DeepSeek model id such as deepseek-chat. The request shape, streaming and response objects stay unchanged.
In DeepSeek's API, what do the deepseek-chat and deepseek-reasoner names select?
basics
~10 sdeepseek-chat selects DeepSeek's general-purpose V3 chat line; deepseek-reasoner selects the R1-style reasoning line that thinks before answering. Both are moving aliases pointing at whatever generation is current, not frozen model versions.
In DeepSeek's API, what is reasoning_content in a deepseek-reasoner response?
basics
~20 sdeepseek-reasoner puts its chain of thought in message.reasoning_content and the user-facing answer in message.content. Both come back in one ordinary chat-completion response. Show users content; treat reasoning_content as diagnostic text you log, collapse, or discard.
What does DeepSeek's response_format json_object guarantee, and what does it not?
basics
~20 sDeepSeek's JSON output mode constrains the model to emit syntactically valid JSON, nothing more. It does not validate against a schema you supply, so field names, types and required keys remain your responsibility to prompt for and then verify.
In DeepSeek's API response, what do prompt_cache_hit_tokens and prompt_cache_miss_tokens mean?
basics
~20 sDeepSeek splits a request's input tokens into two billed buckets: prompt_cache_hit_tokens, served from its context cache at a steeply discounted input rate, and prompt_cache_miss_tokens, billed at the full input rate. The two always sum to prompt_tokens.
What does tokenizer.apply_chat_template() do when prompting a Llama model?
basics
~20 sapply_chat_template() turns a list of {role, content} messages into the exact prompt string the checkpoint was fine-tuned on, using the Jinja template shipped alongside that model's tokenizer — so the special tokens come from the model, not from your code.
Why is Meta's Llama called open-weight rather than open-source?
basics
~20 sLlama weights are downloadable and usable commercially, but the Llama Community License imposes field-of-use limits through an acceptable-use policy and singles out very large operators for extra permission. Those restrictions fail the Open Source Definition, so the accurate term is open-weight.
How much memory do Llama 70B weights need at FP16 versus 4-bit quantization?
basics
~20 sWeights cost bytes-per-parameter times parameter count: about 140 GB for 70B at FP16, about 70 GB at 8-bit, and roughly 35-40 GB at 4-bit once block scales are counted. KV cache and activations are extra and are not shrunk by weight quantization.
When should you serve Llama with Ollama versus vLLM?
basics
~20 sOllama wraps llama.cpp for local, single-user use: one command pulls a quantized model and runs it on a laptop, CPU included. vLLM is a GPU server built for concurrent production traffic, using continuous batching and a paged KV cache.
How does Llama 3's chat template differ from Llama 2's [INST] and <<SYS>> format?
basics
~20 sLlama 2 wraps each user turn in [INST] ... [/INST] plain-text markers and nests the system prompt inside the first one with <<SYS>>. Llama 3 drops both, using real special tokens: a role header per turn, closed by <|eot_id|>.
What sizes does Gemma 3 come in, and which of them accept image input?
basics
~20 sGemma 3 ships at 1B, 4B, 12B and 27B parameters, each in a pretrained and an instruction-tuned variant. The 4B, 12B and 27B checkpoints accept images alongside text; the 1B is text-only and has a much shorter context window.
How do the Gemma Terms of Use differ from an OSI licence like Apache 2.0?
basics
~20 sGemma weights ship under Google's own Gemma Terms of Use plus a prohibited-use policy, not an OSI-approved open-source licence. Commercial use is free with no user or revenue threshold, but you accept use restrictions and must pass those terms to anyone you distribute a derivative to.
Why can a fine-tuned Gemma 3 degrade at inference if the chat template is applied wrong?
basics
~20 sGemma is trained on an exact turn format — start_of_turn markers with user and model roles, closed by end_of_turn, and a single leading BOS token. Serving prompts that differ from the training format, or that duplicate BOS, puts the model off-distribution and quality drops silently.
Which Gemma 3 size fits a single 24 GB GPU, and what do you trade away?
basics
~20 sAt bfloat16 you need roughly 2 GB per billion parameters, so 24 GB holds 4B or 12B but not 27B's ~54 GB. Google's quantisation-aware-trained int4 checkpoints shrink 27B to roughly 15 GB, which fits — at the cost of some quality and nearly all your KV-cache headroom.
How should Gemma's licence terms shape choosing it over an Apache-2.0 open-weight model?
basics
~20 sDecide by how you deliver the model. If you serve it yourself, Gemma's terms cost you almost nothing. If you ship weights inside a product, its use restrictions and pass-along duties follow every copy, while an Apache-2.0 model imposes only attribution and carries a patent grant.
Which base_url and API key let the OpenAI SDK call Alibaba's DashScope Qwen models?
basics
~10 sPoint the OpenAI client at Model Studio's compatible-mode base URL — https://dashscope-intl.aliyuncs.com/compatible-mode/v1 for Singapore, https://dashscope.aliyuncs.com/compatible-mode/v1 for Beijing — and pass your DASHSCOPE_API_KEY as api_key. Model ids are Qwen names such as qwen-plus.
Which Qwen variant do you pick for coding, vision or embedding work?
basics
~10 sQwen ships task-specific lines beside its general chat models: Qwen3-Coder for code and agentic coding, Qwen3-VL for images and video, Qwen3-Embedding and Qwen3-Reranker for retrieval, and the Omni/Audio line for speech input.
How do you serve Qwen weights behind an OpenAI-compatible endpoint with vLLM?
basics
~20 svllm serve Qwen/Qwen3-8B pulls the weights from Hugging Face and starts an OpenAI-shaped HTTP server on port 8000. Point any OpenAI client at http://localhost:8000/v1, pass any non-empty API key, and set model to the served name.
In the Qwen model name Qwen3-30B-A3B-Instruct-2507, what does each part mean?
basics
~20 sQwen3 is the generation, 30B the total parameter count, A3B the roughly 3B parameters active per token in a Mixture-of-Experts layout, Instruct the post-trained chat checkpoint (versus Base or Thinking), and 2507 the July 2025 refresh of that checkpoint.
What does Qwen's ChatML chat template wrap around each message, and why does it matter?
basics
~20 sQwen uses ChatML: every message is rendered as <|im_start|>role, a newline, the content, then <|im_end|>. Generation starts from a trailing <|im_start|>assistant header, and <|im_end|> is the stop token. Wrong formatting means the model rambles or never stops.
How do you name a Claude Opus model in Anthropic's API, and why pin the dated snapshot?
basics
~20 sAnthropic publishes each Claude Opus release as a dated snapshot ID such as claude-opus-4-1-20250805, plus a floating alias such as claude-opus-4-1 that resolves to the newest snapshot in that line. Production should pin the dated snapshot.
How does the Claude Opus model ID differ across Anthropic's API, Bedrock, and Vertex AI?
basics
~10 sThe same Opus release carries a different identifier on each platform: claude-opus-4-1-20250805 on Anthropic's own API, anthropic.claude-opus-4-1-20250805-v1:0 on Amazon Bedrock, and claude-opus-4-1@20250805 on Google Vertex AI.
Anthropic retires the Claude Opus snapshot your service pins — how do you run the migration?
basics
~20 sTreat it as a dependency upgrade with a hard deadline: inventory every place the retired Opus snapshot ID appears, run a golden-set eval against the successor snapshot, fix prompt and parsing regressions, canary the new pin, and finish before the shutdown date, because retired IDs return errors rather than falling back.
In the Anthropic API, what model string identifies Claude Sonnet, and does it take a date suffix?
basics
~10 sCurrent Claude Sonnet models are called by undated alias strings such as claude-sonnet-5 and claude-sonnet-4-6. The ID is complete as written; appending a -YYYYMMDD snapshot suffix produces an unknown model and the request fails.
Which sampling parameters does claude-sonnet-5 reject, and what controls reasoning depth instead?
basics
~10 sclaude-sonnet-5 removed temperature, top_p and top_k — sending any of them returns a 400. The fixed thinking budget (budget_tokens) is gone too. Depth and token spend are set with output_config.effort, alongside adaptive thinking.
After moving a streaming app to claude-sonnet-5, thinking blocks arrive empty — why?
basics
~20 sOn claude-sonnet-5 the thinking parameter's display field defaults to omitted, so thinking blocks stream with empty text. The previous generation defaulted to summarized. Set thinking to {"type": "adaptive", "display": "summarized"} to get readable reasoning back.
In OpenAI's Images API, how does a gpt-image generation return the image data?
basics
~20 sgpt-image models always return base64-encoded bytes in the response's data[].b64_json field. There is no hosted URL to fetch and no response_format choice, so the caller decodes the string and stores or serves the bytes itself.
Migrating OpenAI image calls from DALL-E 3 to gpt-image: which parameters change?
basics
~20 sThe endpoint path stays the same, but the parameter vocabulary changes: new size values, a low/medium/high quality scale instead of standard/hd, no style and no response_format, no revised_prompt in the response, and new background, output_format and moderation options.
In OpenAI's /v1/images/edits endpoint, what must the mask file look like?
basics
~20 sThe mask is an optional PNG with an alpha channel, uploaded alongside the image and matching its dimensions. Fully transparent pixels mark the region the model may repaint; opaque pixels are preserved. Omit it and the whole image is re-rendered.
How do you handle a moderation refusal from OpenAI's Images API in production?
basics
~20 sA blocked image request returns HTTP 400 with a moderation_blocked error, not a rate-limit or server error. It is not retryable: surface a clear refusal, invite the user to rephrase, log the attempt for abuse review, and never loop on backoff.
How does OpenAI bill a gpt-image generation, and what drives the cost?
basics
~20 sgpt-image calls are billed in tokens, not at a flat per-image price. Text prompt tokens, any input image tokens, and image output tokens are counted separately; output tokens scale with the requested size and quality tier.
OpenAI's audio transcription endpoint caps uploads at 25 MB — how do you transcribe a two-hour recording?
basics
~20 sRe-encode to mono 16 kHz compressed audio first, since Whisper resamples to 16 kHz anyway. If the file is still over the limit, split it at silence boundaries, transcribe each chunk, then add each chunk's start offset to its timestamps before merging.
In OpenAI's audio API, how do /v1/audio/transcriptions and /v1/audio/translations differ?
basics
~20 sTranscription writes down what was said in the language it was spoken in. Translation always outputs English, whatever the source language — there is no target-language parameter, because the model was only trained to translate into English.
What does the prompt parameter do on an OpenAI Whisper transcription request?
basics
~20 sIt supplies prior text as decoder context, biasing spellings, jargon and punctuation style toward what the prompt contains. It is not an instruction field — Whisper does not follow commands in it — and only roughly the last 224 tokens are used.
How do you get word-level timestamps from OpenAI's Whisper transcription API?
basics
~20 sSend timestamp_granularities: ["word"] on the transcription request and set response_format to verbose_json; any other response format rejects the granularity option. The response then carries a words array with a start and end time per word, alongside the usual segment list.
Whisper transcripts contain repeated phrases over silent audio — how do you diagnose and reduce it?
basics
~20 sThis is Whisper's known hallucination on non-speech: with no acoustic evidence the decoder falls back on language priors and loops or emits training-set boilerplate. Strip silence with voice-activity detection before transcribing, and drop segments whose no_speech_prob is high or compression_ratio suggests repetition.
How does OpenRouter bill a call, and where do its per-model prices come from?
basics
~10 sOpenRouter runs on prepaid credits: you top up one balance and each call deducts the chosen provider's own per-token price. GET /api/v1/models publishes every model's pricing object, quoted in US dollars per single token.
How do you point the OpenAI SDK at OpenRouter, and how are models named there?
basics
~10 sSet the OpenAI client's base URL to https://openrouter.ai/api/v1 and send your OpenRouter key as a bearer token. Models are addressed by a vendor-prefixed slug such as anthropic/claude-sonnet-4.5, so one key reaches every vendor.
How do you make OpenRouter return the actual cost and native token counts inline?
basics
~20 sSend "usage": {"include": true} in the chat completion request body. The response's usage object then carries cost in credits alongside token counts from the serving model's own tokenizer, with cost_details breaking out any upstream charge.
In OpenRouter, what does the models array in a chat completions request do?
basics
~20 sOpenRouter treats models as an ordered fallback list: it tries the first entry and, if that model errors or is unavailable, retries the next one inside the same request. You are billed for whichever model actually answered.
How do OpenRouter's provider.order and allow_fallbacks fields control routing?
basics
~20 sprovider.order lists upstream providers to try in priority order for the chosen model. allow_fallbacks, true by default, decides what happens when none of them can serve: leave it on and OpenRouter tries other providers, set it false and the request fails instead.
What is GitHub Models, and when would you use it instead of a paid provider account?
basics
~20 sGitHub Models is a hosted catalog and playground of models you can call with your GitHub token. It gives free, rate-limited inference for prototyping, so you can compare models from several publishers before opening any paid provider account.
How do you authenticate to the GitHub Models inference API from code and from CI?
basics
~20 sSend a GitHub token as an HTTP bearer credential. Locally that is a fine-grained personal access token carrying the Models read permission; inside GitHub Actions it is the workflow's built-in GITHUB_TOKEN, after the job declares the models: read permission.
How do you run GitHub Models prompt evaluations from a GitHub Actions workflow?
basics
~20 sStore prompts as .prompt.yml files in the repository, grant the job models: read so the workflow token can call inference, install the gh-models CLI extension, and run gh models eval on those files so prompt changes are reviewed and scored like code.
Which SDKs call GitHub Models, and what changes in OpenAI SDK code to target it?
basics
~20 sTwo client styles work: the Azure AI Inference SDK, and any OpenAI-compatible client pointed at the GitHub Models base URL. With the OpenAI SDK you override base_url, pass the GitHub token as api_key, and use publisher-qualified model ids such as openai/gpt-4o-mini.
Your GitHub Models prototype starts returning HTTP 429 under load — how do you respond?
basics
~20 sTreat 429 as quota, not a bug: GitHub Models throttles per account on requests per minute and per day, tokens per request and concurrency, with tighter limits on larger models. Back off using the wait the error states, and move sustained load to a paid endpoint.