OpenAI
OpenAI is the reference API of the whole space — the request shape most other vendors ended up imitating. Interviews walk the surface: chat completions, tool calling, structured outputs, embeddings, and the operational side of rate-limit tiers and cost.
on this pageshowhide
explore
- Chat Completions API6 questions
- Tool & Function Calling6 questions
- Structured Outputs5 questions
- Embeddings API6 questions
- Assistants & Responses API6 questions
- Rate Limits & Cost6 questions
questions
page 2 of 2How do you migrate a production vector index to a new OpenAI embedding model?
basics
~20 sVectors from different embedding models are not comparable, so the whole corpus must be re-embedded. Build the new index alongside the old one, keep queries on the old index until the new one is complete and evaluated, then cut over and only afterwards delete the old vectors.
How do you forecast and cap OpenAI spend before shipping a new feature?
basics
~20 sBuild a unit-cost model from a pilot: measured p50 and p95 input and output tokens per request, multiplied by the model's per-million rates and forecast volume. Then enforce it with per-project budgets, separate keys per feature for attribution, and alerts on tokens per request.
When would you not put a strict JSON schema on an OpenAI call, and what do you use instead?
basics
~20 sSkip strict schemas when the output is prose, when the shape is decided per request so every call pays fresh grammar compilation, when the schema cannot fit the restricted dialect or its size caps, or when portability across providers outweighs the guarantee.
What are the limits of OpenAI's hosted code_interpreter tool container?
basics
~20 sThe tool runs model-written Python in a sandboxed container with no open internet access. The container is short-lived, expiring after a period of inactivity, is billed per session on top of tokens, and any files it produces are lost unless you download them first.
How many tools should one OpenAI request expose, and what breaks with too many?
basics
~20 sTool definitions are serialized into the prompt and billed as input tokens on every request, and selection accuracy falls as similar tools multiply. Keep the per-request set small and distinct; route or split across sub-agents rather than growing one surface.
showing 31–35 of 35