skip to content

What does adding tf.lite.OpsSet.SELECT_TF_OPS to a TFLite conversion cost you?

level: seniorimportance: should knowfreq 46%

answer

  1. small built-in kernel set is the whole point
  2. escape hatch, not a free upgrade
  3. app binary grows by tens of megabytes
  4. those ops are CPU-only
  5. graph splits, copies at every seam

basics

~20 s

It unblocks conversion by allowing selected TensorFlow kernels into the model, at the price of a much larger app binary, CPU-only execution for those ops, a fragmented graph that accelerators cannot take whole, and ops that generally will not quantize.

solid answer

~50 s

TFLite ships a deliberately small built-in kernel library. When your model uses an op outside it, `convert()` fails naming that op. Setting `converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS, tf.lite.OpsSet.SELECT_TF_OPS]` lets the converter emit those ops as TensorFlow ops instead, and conversion succeeds. The bill arrives on device: the app must now link the Select TF ops runtime, which adds tens of megabytes compared with the built-in-only runtime; those ops execute on CPU, so a graph that would have run entirely on an accelerator is cut into pieces with copies at each seam; and they usually stay float, which undermines a full-integer conversion. Treat it as a bridge, not a destination — the better fixes are usually rewriting the offending layer in supported ops, moving preprocessing out of the graph, or pinning the shapes that forced a dynamic op in the first place.

code

python · 8 lines
python
import tensorflow as tf

converter = tf.lite.TFLiteConverter.from_saved_model("saved_model/")
converter.target_spec.supported_ops = [
    tf.lite.OpsSet.TFLITE_BUILTINS,   # prefer built-in kernels
    tf.lite.OpsSet.SELECT_TF_OPS,     # fall back to TF kernels
]
open("model_flex.tflite", "wb").write(converter.convert())

go deeper

for a junior

Recognize that conversion can fail because an operator has no TFLite implementation, and that there is a flag which allows TensorFlow operators through. Know it is not free.

for a middle

Name the concrete costs — a much larger runtime dependency, CPU-only execution, and no int8 path for those ops — and show the correct two-element supported_ops list rather than replacing the built-ins.

for a senior

Lead with the alternatives: move preprocessing out of the exported signature, pin shapes, or rewrite the layer. If you do enable the fallback, describe measuring binary size and end-to-end on-device latency before and after.

for a principal

Own it as a release-level tradeoff: app download size against engineering time, whether a selective runtime build belongs in the pipeline, and who is accountable for removing the fallback rather than letting it become permanent.

## Why conversion refuses in the first place TensorFlow has thousands of operators. TFLite implements a few hundred, chosen because they are what inference models actually need and because every kernel is code that must be compiled into a mobile app. When the converter hits an op with no built-in equivalent and no lowering, it raises an error naming the op and pointing at the Select TF ops escape hatch. The ops that trigger this cluster in predictable places: text and string processing, ragged or sparse tensors, data-pipeline ops that leaked into the serving signature, control flow with dynamic trip counts, some RNG ops, and newer or exotic layers that have not been given a TFLite kernel. ## What the flag does ``` converter.target_spec.supported_ops = [ tf.lite.OpsSet.TFLITE_BUILTINS, tf.lite.OpsSet.SELECT_TF_OPS, ] ``` List both. `TFLITE_BUILTINS` keeps the built-in kernels preferred, and `SELECT_TF_OPS` permits the remainder to be emitted as TensorFlow ops embedded in the FlatBuffer. Conversion now succeeds. ## The four costs **Binary size.** Those ops are not implemented in the TFLite runtime, so the app has to ship the Select TF ops runtime library that hosts real TensorFlow kernels. That is a large dependency — tens of megabytes of added download size in the default configuration, against a built-in-only runtime measured in single-digit megabytes. On consumer mobile apps, download size is a business metric, and this alone kills the option in many teams. Selective builds that include only the kernels your model uses can claw much of it back, at the cost of a custom build step in your release pipeline. **CPU only.** The Select TF ops kernels run on the CPU. Any accelerator you were counting on cannot execute them. **Graph fragmentation.** This is the one candidates miss. If your unsupported op sits in the middle of the network, the graph splits into a supported prefix, a CPU-only middle, and a supported suffix. Each seam is a synchronization point and a tensor copy between the accelerator's memory and the CPU's. A model that was 5 ms end to end on an accelerator can become dramatically slower once it is cut in three, and the cause is not the op's own cost but the round trips around it. **Quantization.** Select TF ops generally have no int8 path, so they stay float. If you were converting with a representative dataset targeting an integer-only device, the presence of a float island can mean the model is rejected outright. ## What to do instead Before reaching for the flag, ask why the op is there. *Is it preprocessing?* String splitting, tokenization, image decoding and normalization frequently end up inside the exported signature because they were part of the training graph. Move them into application code. This is the single most common real fix and it costs nothing at run time. *Is it a shape problem?* Ops that only appear because a dimension is dynamic often vanish once you export with a fixed `input_signature`, letting the converter constant-fold the surrounding logic. *Is it one layer?* Rewriting a custom layer in terms of supported primitives is often a half-day of work and permanently removes the dependency. *Is it genuinely unavoidable?* Then you have two paths. `SELECT_TF_OPS` gets you running today. Alternatively `converter.allow_custom_ops = True` lets the op through as a custom op with no kernel attached, and you supply a hand-written implementation registered with the runtime — far more work, but it keeps the small runtime and lets you write an accelerator-friendly kernel. ## How to talk about it in an interview The weak answer is "add SELECT_TF_OPS and it works." The strong answer names the tradeoff and then says what you would measure: which ops actually forced it (the converter error tells you), whether they are in the hot path or at the edges, what the binary size delta is against your app's budget, and what the on-device latency looks like with and without. Then it presents the flag as a deliberate, temporary choice with an owner and a follow-up, rather than a fix.

  • Which ops most often force a team to reach for SELECT_TF_OPS?
    String and text ops, ragged or sparse tensor ops, data-pipeline ops that leaked into the exported signature, dynamic control flow, and some RNG ops. The pattern is telling: much of it is preprocessing that was never meant to live inside the inference graph. Moving that work into application code removes the need for the flag entirely.
  • How does converter.allow_custom_ops differ from SELECT_TF_OPS?
    SELECT_TF_OPS embeds real TensorFlow kernels and requires shipping the Select TF ops runtime. allow_custom_ops lets an unrecognized op through as an opaque custom op with no implementation in the file — you must write and register a kernel in the app yourself. It is much more work, but it keeps the small runtime and lets you write a kernel tuned for your target.
  • Why can a model with one Select TF op be far slower than the op itself costs?
    Because it partitions the graph. Everything before the op runs on the accelerator, the op runs on CPU, and everything after goes back to the accelerator, with a tensor copy and a synchronization at each boundary. On a small model those round trips can dominate total latency, so the honest measurement is end-to-end on device, not the op's isolated cost.

saying these in an interview costs you the question

  • Treats SELECT_TF_OPS as a free compatibility switch
  • Forgets the app must ship the extra runtime library
  • Assumes Select TF ops can run on an accelerator
  • Expects those ops to quantize to int8 like builtins
  • Omits TFLITE_BUILTINS and keeps only SELECT_TF_OPS

context