skip to content

How would you design OSPF link costs for a network mixing 100 Gb/s core links, 10 Gb/s access links and slow WAN circuits?

level: principalimportance: nice to knowfreq 8%

answer

  1. costs encode design, not load
  2. headroom above the fastest link
  3. the 16-bit ceiling for slow links
  4. bandwidth ignores latency
  5. one policy, enforced everywhere

basics

~10 s

Choose one deliberate OSPF cost scheme and apply it everywhere: a reference bandwidth above the fastest link that keeps the slowest inside the 16-bit cost field, or role-based costs that encode latency and intent.

solid answer

~50 s

There is no single right scheme, so I start from what costs must express. A **bandwidth-derived** scheme is cheap to run: set the reference above today's fastest link with headroom — with 400 Gb/s, 100G costs 4, 10G costs 40 and 10 Mb/s costs 40,000, inside the 65,535 limit — on every router. Its weaknesses: it ignores latency, so a long 10G path beats a short slower one, and where an implementation derives cost from a bundle's current bandwidth, losing a member changes the cost and triggers flooding. A **role-based** scheme — explicit costs per tier such as core, access and WAN — encodes intent and latency and keeps parallel paths equal, but needs documentation and review. Either way costs are static, consistency must be enforced by configuration management, and when integers can't steer traffic, that points to traffic engineering.

go deeper

for a junior

Remember that OSPF cost is a number the network team chooses, and the choice decides which paths traffic prefers.

for a middle

Be ready to compute costs for a given reference bandwidth and to say why a reference below the fastest link makes fast links look identical.

for a senior

Spot the failure modes of each scheme: clamped slow links, latency-blind paths, bundle churn, and inconsistent references causing asymmetric routing.

for a principal

Own the policy: pick bandwidth, role-based or hybrid costs for stated goals, plan headroom for faster links, and make consistency enforceable rather than hoped for.

## What an OSPF cost can and cannot express RFC 2328 gives each router interface a single, dimensionless **output cost**, chosen by the operator, greater than zero and carried in a **16-bit** field of the router-LSA (maximum 65,535). Routers sum costs along each path and take the lowest. That is the whole toolbox: - cost is **static** — OSPF never measures utilisation or delay; - cost is **directed** — each end of a link advertises its own output cost; - cost is **summed**, so many cheap hops can lose to one expensive hop or beat it. Any policy is therefore a statement of *intent* frozen into integers. The design question is which intent to encode and how to keep it consistent across hundreds of devices. ## Option 1: bandwidth-derived costs Implementations commonly compute `cost = reference bandwidth / interface bandwidth`, with a default reference (often 100 Mb/s) that makes every modern link cost 1. Raising the reference fixes that: | Link | Reference 100 Gb/s | Reference 400 Gb/s | |---|---|---| | 400 Gb/s | 1 (minimum) | 1 | | 100 Gb/s | 1 | 4 | | 10 Gb/s | 10 | 40 | | 1 Gb/s | 100 | 400 | | 10 Mb/s | 10,000 | 40,000 | Strengths and weaknesses: - **Low effort**: one setting per router, new links get sensible costs automatically. - **Headroom matters**: with a 100 Gb/s reference a later 400 Gb/s link collapses to the same cost as 100 Gb/s. - **The ceiling bites at the slow end**: with 400 Gb/s, a 1 Mb/s circuit would compute to 400,000, which does not fit and gets clamped. - **Latency is invisible**: a 10G path across a continent beats a 1G metro path, even for latency-sensitive traffic. - **Bundles can churn**: where an implementation derives cost from a bundle's *current* bandwidth, losing one member changes the cost, forcing a new router-LSA, flooding and SPF across the area. ## Option 2: role-based manual costs Assign explicit costs by the link's **role** in the design rather than its speed: | Tier | Example cost | Intent | |---|---|---| | Core-to-core | 10 | preferred backbone | | Core-to-access | 100 | normal | | Inter-site WAN | 1,000 | use only when needed | | Backup WAN | 5,000 | last resort | - It **encodes latency and preference** that bandwidth cannot. - It makes **parallel paths deliberately equal** so equal-cost sharing works where you want it. - It needs a **written standard** and review, because every new link is a decision. ## Hybrid Many networks combine them: bandwidth-derived inside a site, where latency is uniform, with **manual overrides** on WAN and backup links. The risk is drift — overrides nobody remembers. ## Operational guardrails 1. **One policy, everywhere.** A router with a different reference or rule advertises different costs, skewing paths and making the two directions of a flow asymmetric. 2. **Template it.** Generate interface costs from configuration management, not hand edits. 3. **Check for asymmetry**: compare the two ends of every link and flag mismatches. 4. **Simulate before changing**: a cost change shifts traffic across the whole area at once. ## Deciding | If the network… | Lean toward | |---|---| | is one site with uniform latency | bandwidth-derived with headroom | | spans regions or mixes satellite, metro and long-haul links | role-based or hybrid | | needs load-aware or per-flow steering | costs for the baseline, traffic engineering for the rest | ## Common traps - **Changing the reference on a few routers at a time.** During a rollout the network runs mixed references, so stage it in a window and finish it quickly. - **Encoding temporary drains into the permanent scheme.** A cost raised for maintenance and never restored quietly becomes policy. - **Tuning by trial and error.** Each cost tweak can move traffic area-wide; a model of the topology predicts the effect before it reaches production. What you are being judged on is not the integers but the reasoning: what the costs must express, how the scheme survives growth to faster links, and how it stays consistent when many people touch the network.

  • Why can bandwidth-derived OSPF costs on a link bundle cause routing churn?
    Where an implementation computes cost from the bundle's current bandwidth, losing one member lowers that bandwidth, raises the cost, and forces a new router-LSA — flooding and SPF across the area for a failure that may not need any rerouting. Designs often pin the bundle's cost by hand so a member failure stays local. How the bundle itself behaves is a link-aggregation subject.
  • Could OSPF costs be made to follow live congestion?
    Not in the protocol: RFC 2328 cost is one static, dimensionless value per output interface. Making it dynamic would mean reoriginating LSAs as utilisation moves, flooding and rerunning SPF constantly, and risking oscillation as traffic flips between paths and back. Load-aware steering belongs to traffic engineering or capacity planning, with OSPF costs describing the baseline design.

saying these in an interview costs you the question

  • RFC 2328 says the reference bandwidth should match the fastest link.
  • OSPF costs should follow live utilisation to balance load.
  • Any reference bandwidth works; raising it further has no downside.
  • Bandwidth-derived costs already account for link latency.
  • Manual costs are always better than a bandwidth formula.