Where do multipart temp files live, how are they cleaned up, and how would you design a memory-safe path for very large uploads?
answer
- spool > file-size-threshold to location temp dir
- container owns + deletes temp files at request end
- consume within request; transferTo may move not copy
- stream getInputStream, never getBytes for big files
- large/resumable -> presigned direct-to-S3
basics
~20 sParts above file-size-threshold spool to the temp dir (spring.servlet.multipart.location, default container temp). The container deletes them when the request ends. For huge uploads, stream via getInputStream()/transferTo() to storage instead of getBytes(), and cap sizes to bound resources.
solid answer
~50 sWith StandardServletMultipartResolver the container spools parts larger than file-size-threshold to disk under spring.servlet.multipart.location (default the servlet temp dir), keeping smaller parts in memory. These temp files are the container's responsibility and are deleted when the request completes, so you must read/transfer within the request thread — transferTo() is efficient because it can move the already-spooled file. For very large uploads: never use getBytes() (full heap copy); stream getInputStream() to the destination (disk/S3) with bounded buffers; set max-file-size/max-request-size and align them with container limits to prevent DoS; put location on a volume with capacity, correct permissions, and consider isolation from other tenants. Guard the filename (path traversal), validate real content type (magic bytes, not the client header), and consider antivirus/zip-bomb checks. For truly large or resumable uploads, prefer chunked/resumable protocols or direct-to-object-store presigned uploads over classic multipart.
code
java · 17 lines@PostMapping(path = "/large", consumes = MediaType.MULTIPART_FORM_DATA_VALUE)
public ResponseEntity<String> uploadLarge(@RequestParam("file") MultipartFile file) throws IOException {
// Server-generated name -> no path traversal from client filename
String key = UUID.randomUUID().toString();
// Stream to storage with bounded memory; never file.getBytes() for big files.
try (InputStream in = new BufferedInputStream(file.getInputStream())) {
storageClient.put(key, in, file.getSize()); // e.g. S3 putObject with content length
}
// Or, for local disk, transferTo can MOVE the already-spooled temp file:
// file.transferTo(uploadRoot.resolve(key).normalize());
return ResponseEntity.ok(key);
}
// Note: the container's temp file is cleaned up automatically when the request completes,
// so all reading/transfer must happen inside this handler thread — not on a later async task.go deeper
Know temp files exist in a temp dir and Spring/container cleans them up.
Explain file-size-threshold spooling, transferTo efficiency, and streaming vs getBytes.
Cover temp-dir capacity/permissions, request-scoped cleanup constraint, and size/path/type guards.
Architect a DoS- and abuse-resistant upload path, and know when to bypass app-server multipart entirely via presigned/resumable uploads.
## Temp-file lifecycle - With `StandardServletMultipartResolver`, the **servlet container** does the spooling. A part is buffered in memory until it exceeds **`file-size-threshold`** (default `0B` = spool immediately), then written to a temp file in **`spring.servlet.multipart.location`** (default the container's temp dir, e.g. `${java.io.tmpdir}` / Tomcat work dir). - The temp files are owned by the container. On request completion, `DispatcherServlet.cleanupMultipart()` runs and the container **deletes** the spooled parts. Consequence: you must consume the `MultipartFile` **during the request** — stashing it for a background thread yields a missing/closed file. - **`transferTo(dest)`** is efficient precisely because the part is often already a temp file on disk: the implementation can *move/rename* it rather than re-copy, and it also relieves you of cleanup. ## Memory safety for large uploads - **Avoid `getBytes()`** — it materializes the whole file as a heap `byte[]`; N concurrent large uploads multiply heap use and invite `OutOfMemoryError`. - **Stream** with `getInputStream()` (bounded buffer) to the sink, or `transferTo()` straight to final storage. - Keep **`file-size-threshold` low** (default disk) so large parts don't sit in heap. - **Cap sizes** via `max-file-size` / `max-request-size`, and **align container limits** (Tomcat `maxSwallowSize`, connector caps) so oversized requests are rejected cleanly rather than resetting connections. - Watch **back-pressure**: multipart is buffered fully (parsed) before the handler with eager parsing, so the disk/temp volume must absorb the peak. `resolve-lazily=true` defers parsing but complicates error handling. ## Temp-dir operational concerns - **Capacity**: a full `location` volume aborts uploads mid-stream with IOExceptions. Size it for peak concurrency × max-request-size. - **Permissions & isolation**: the dir must be writable by the app user; avoid a world-readable shared `/tmp` where other processes could read sensitive uploads. Prefer a dedicated, restricted directory. - **Cleanup on crash**: if the JVM/container dies mid-request, temp files may linger — have a janitor or rely on the container's startup cleanup. ## Security hardening (principal-level) - **Path traversal**: never build the save path from `getOriginalFilename()`; generate a server-side name and `normalize()`/verify the resolved path stays under the upload root. - **Content-type spoofing**: `getContentType()` is client-declared. Validate the *actual* type (magic-byte sniffing, e.g. Apache Tika) if type matters for safety. - **Malware / zip bombs**: scan with AV; bound decompression ratios; don't auto-expand archives. - **Storage location**: never write user uploads into a web-served/static path where they could be executed or served with an attacker-controlled type. ## When classic multipart isn't the right tool - **Very large / resumable** uploads (GBs, flaky networks): a single multipart POST is fragile. Prefer **chunked/resumable** protocols (tus, multipart chunking) or **presigned direct-to-object-store** uploads (S3/GCS) so bytes never transit or buffer through the app server — the app only issues the signed URL and records metadata. - **Throughput/scaling**: offloading to object storage removes the app server as a disk/memory bottleneck and simplifies horizontal scaling. ## Summary decision guide - Small/medium files, form-driven → standard multipart + streaming `transferTo`, tuned limits, name/type guards. - Large/resumable/high-scale → presigned direct uploads or resumable protocol; keep the app out of the byte path.
- Why can't you hand a MultipartFile to a background thread to save later?The container-spooled temp file backing it is deleted during multipart cleanup when the request completes. By the time the async task runs, the underlying file is gone or the stream is closed. Consume it within the request, or copy to your own persistent location first.
- For multi-GB or unreliable uploads, why prefer presigned direct-to-object-store over a multipart POST?A single multipart POST has no resumability, buffers/streams the full payload through the app server (disk/memory/bandwidth bottleneck), and fails wholesale on a dropped connection. Presigned uploads send bytes straight to S3/GCS with client-side retries/chunking; the app only mints the URL and stores metadata, so it scales and stays out of the byte path.
- Why is getContentType() insufficient for validating file type?It's the client-declared MIME type in the part header and is trivially spoofable. For safety-relevant checks, sniff the real type from magic bytes (e.g., Apache Tika) rather than trusting the header.
saying these in an interview costs you the question
- Storing the MultipartFile reference and saving it from an async job after the request ends.
- Using getBytes() for multi-hundred-MB uploads.
- Trusting getContentType()/getOriginalFilename() for security decisions.
- Writing uploads into a web-served static directory.
- Assuming a single multipart POST is fine for GB-scale resumable uploads.