What are the serialization, sizing, and consistency pitfalls of the ExecutionContext, and how do they influence how you use it?
answer
- Jackson serializer default; trusted-type / whitelist concern
- SHORT_CONTEXT length-limited -> keep tiny
- step ctx committed with chunk -> consistent, rolls back
- namespace keys (setName) to avoid collisions
- control channel, not data bus
basics
~20 sThe context is serialized into size-limited DB columns, so keep it small and serializable. It's saved per chunk with the transaction, so it stays consistent with the data. Don't store large objects or non-serializable types, and namespace keys to avoid collisions.
solid answer
~40 sThe ExecutionContext is serialized by an ExecutionContextSerializer (Jackson-based by default in modern Spring Batch) and stored in the BATCH_*_EXECUTION_CONTEXT tables, with a SHORT_CONTEXT column that is length-limited. Design implications: keep entries tiny (scalars, small maps) and serializable by the configured serializer; large or non-serializable values cause failures or bloat. Because the step context is persisted transactionally at each chunk commit, progress stays consistent with committed data — a rollback also reverts the context. Watch for key collisions in a shared step context: namespace reader keys via setName. With the Jackson serializer, custom types may need to be trusted/whitelisted to deserialize. Treat the context as a small, durable control channel for resumability, not a data-transfer mechanism — bulk data belongs in staging tables or the item pipeline.
code
java · 13 lines// Overriding the ExecutionContext serializer on the JobRepository.
@Bean
public JobRepositoryFactoryBean jobRepository(DataSource ds,
PlatformTransactionManager tx) {
JobRepositoryFactoryBean f = new JobRepositoryFactoryBean();
f.setDataSource(ds);
f.setTransactionManager(tx);
Jackson2ExecutionContextStringSerializer serializer =
new Jackson2ExecutionContextStringSerializer();
// Only tiny, serializable values should ever go into the context.
f.setSerializer(serializer);
return f;
}go deeper
Unlikely to know serialization/sizing details.
May know 'keep it small' but not the transactional or serializer specifics.
Should explain size limits, transactional coupling, and key hygiene.
Should reason about serializer choice, trusted-type security, restart determinism, and the context as a design boundary.
## Serialization The `JobRepository` writes the context through an `ExecutionContextSerializer`. Modern Spring Batch defaults to a **Jackson-based** serializer (`Jackson2ExecutionContextStringSerializer`), replacing the old XStream/JDK approaches. Consequences: - Values must be **serializable/deserializable** by that serializer. Plain scalars and simple beans are fine; exotic or non-Jackson-friendly types can fail on write or read. - For arbitrary custom classes, Jackson deserialization may require the type to be **trusted** (whitelisted) — a security measure so the metadata store can't be used to instantiate arbitrary classes. If you plug in a different serializer, you own its safety. - You can override the serializer via `JobRepositoryFactoryBean.setSerializer(...)` (or the corresponding Boot customization). ## Sizing The schema stores a **`SHORT_CONTEXT`** (a length-bounded VARCHAR — historically ~2500 chars, larger in current schemas) plus a `SERIALIZED_CONTEXT` for overflow depending on the DB. Either way, the context is meant to be **small**. Anti-patterns: - Stuffing collections of domain objects, blobs, or entire result sets. - Accumulating unbounded state (a growing list) that inflates every chunk commit and every restart load. Bloat slows every chunk (each commit re-serializes and writes the context) and can overflow columns. ## Consistency / transactional coupling The **step** context is persisted **within the chunk transaction**. So: - On commit, saved progress exactly reflects committed data. - On rollback, the context update is rolled back too — no drift between 'what we processed' and 'what we recorded'. This is the backbone of correct resume. The **job** context is not per-chunk; it's persisted on job-execution updates, so don't rely on it for mid-step durability. ## Key hygiene A single step's context is shared by all its streams and listeners. Collisions corrupt resume. Built-in readers avoid this by prefixing keys with the reader's **name** (`setName`, via `ExecutionContextUserSupport.getKey(...)`). Always give custom stateful readers a unique name and namespaced keys. ## The `dirty` flag `ExecutionContext` tracks a `dirty` boolean; mutating methods set it, and the framework uses it to decide whether a re-persist is needed. Reading values does not dirty it. ## When to use / not use - **Use** for: read counts, cursors/offsets, small computed scalars to promote across steps, restart bookkeeping. - **Don't use** for: passing large payloads between steps (use staging tables, files, or the reader/writer pipeline), or as an application cache. ## Restart correctness recap Correct resume requires: a restartable job, deterministic input ordering (for count-based fast-forward), transactional writers to avoid duplicates, and small serializable context state. Violating any of these — non-deterministic queries, non-serializable values, oversized context — undermines the guarantees.
- Why must you keep the ExecutionContext small?It is serialized and written to a length-limited column on every chunk commit; large values cause slow commits, column overflow, and bigger restart-load cost.
- Why does the Jackson serializer care about 'trusted' types?Deserializing arbitrary class names from the metadata store is a security risk (gadget/instantiation attacks), so custom types generally must be explicitly trusted/whitelisted.
- What guarantees the saved read-count never disagrees with the written data?The step context is persisted inside the same transaction as the chunk, so commit saves both together and rollback reverts both.
saying these in an interview costs you the question
- Storing large collections or blobs in the context
- Assuming any object serializes/deserializes without regard to the serializer or trusted types
- Believing the job context is flushed per chunk like the step context
- Ignoring key namespacing across multiple readers in one step