skip to content

Metrics and Convergence

OSPF cost is reference bandwidth over interface bandwidth, so fast links all look alike until the reference is raised. Interviewers ask what sets convergence time and how to shorten it.

on this pageshow

questions

5

In OSPF, why do a 10 Gb/s link and a 1 Gb/s link often end up with the same interface cost?

level: middleimportance: must knowfreq 42%

answer

  1. the RFC leaves cost to the operator
  2. reference divided by interface bandwidth
  3. a common 100 Mb/s implementation default
  4. integer costs never fall below 1

basics

~20 s

Many implementations derive an OSPF interface's cost as a reference bandwidth, commonly 100 Mb/s by default, divided by the interface bandwidth, never below 1, so every link of 100 Mb/s or faster costs 1 until the reference is raised.

solid answer

~40 s

RFC 2328 only says an interface's cost is set by the administrator, must be greater than zero and travels as a 16-bit `metric` in the router-LSA. The formula `cost = reference bandwidth / interface bandwidth` is an implementation convenience, and a widespread default reference is 100 Mb/s. Because the result is an integer of at least 1, 100 Mb/s, 1 Gb/s and 10 Gb/s links all cost 1, so OSPF cannot tell them apart: it may split traffic equally over a 10G and a 1G link, or prefer one 1G hop (cost 1) over two 10G hops (cost 2). The fix is to raise the reference above the fastest link on every router — with 100 Gb/s, 10G costs 10 and 1G costs 100 — or to set costs by hand.

go deeper

for a junior

Remember that OSPF prefers the path with the lowest total cost and that cost is often derived from bandwidth, so very fast links can look identical.

for a middle

Walk the formula: reference divided by interface bandwidth, rounded, never below 1, with a common 100 Mb/s default reference that makes everything from 100 Mb/s upward cost 1.

for a senior

Show the operational fix: raise the reference consistently on every router, keep the slowest link inside the 16-bit cost range, and explain why a partial rollout skews paths without looping.

for a principal

Frame cost as a design decision rather than a default: whether bandwidth is even the right signal, and how the organisation keeps one cost policy consistent across every device.

## What the standard actually says OSPF (RFC 2328) attaches a **cost** to the *output side* of every router interface. The router advertises that cost as the `metric` of the link in its **router-LSA**, and every router in the area adds up the costs along each candidate path; the lowest total wins. The specification is deliberately silent on how the number is chosen: - the interface output cost is **configured by the administrator**; - it **must be greater than zero**; - in a router-LSA it is a **16-bit** field, so the largest cost a link can carry is **65,535** (0xffff); - RFC 2328 calls it "a single dimensionless metric" — it is not bits per second, not milliseconds, just a number. Nothing in the RFC mentions bandwidth. ## The bandwidth formula is an implementation convention Because asking operators to hand-pick a cost for every interface is tedious, implementations offer a default derived from the interface's configured bandwidth: `cost = reference bandwidth / interface bandwidth` A widespread **implementation default** for the reference is **100 Mb/s**. That choice made sense when 100 Mb/s was a fast link. The result has to be a positive integer, so implementations round it and never go below 1. Treat both the 100 Mb/s figure and the rounding rule as implementation choices, not protocol rules. ## Why fast links collapse to the same cost | Interface | 100 Mb/s reference | 100 Gb/s reference | |---|---|---| | 100 Gb/s | 0.001 -> **1** | **1** | | 10 Gb/s | 0.01 -> **1** | **10** | | 1 Gb/s | 0.1 -> **1** | **100** | | 100 Mb/s | **1** | **1,000** | | 10 Mb/s | **10** | **10,000** | With the old default, everything from 100 Mb/s upward is cost 1. OSPF literally cannot see the difference between a 10G and a 1G link. ## The two-path trap Take routers R1, R2 and R3. R1 has a direct 1 Gb/s link to R2, and also reaches R2 through R3 over two 10 Gb/s links. With the 100 Mb/s reference: 1. The direct path costs **1** (one 1G hop at cost 1). 2. The path through R3 costs **1 + 1 = 2** (two 10G hops at cost 1 each). 3. SPF picks the lowest total, so traffic takes the **1 Gb/s** link and the 10 Gb/s path sits idle. Raise the reference to 100 Gb/s and the totals become **100** for the direct 1G link and **10 + 10 = 20** for the two 10G hops, so traffic moves to the fast path. A second symptom of the same problem: two parallel links of 10G and 1G between the same pair of routers both cost 1, so OSPF treats them as equal-cost next hops and the 1G link saturates first. (How equal-cost traffic is shared is a routing-table feature in its own right.) ## Fixing it - **Raise the reference bandwidth** above the fastest link you have or expect, so each speed gets a distinct cost. - **Do it on every router.** Each router advertises the cost of *its own* outgoing interfaces; receivers never recompute it. A router left on the old reference keeps advertising 1 on its fast links, so paths skew toward it and the two directions of a flow can take different routes. It does not cause a lasting loop — every router runs SPF over the same database — but it does produce suboptimal and asymmetric routing. - **Stay inside the 16-bit field.** With a 100 Gb/s reference a 10 Mb/s link costs 10,000, which fits; a 1 Mb/s link would compute to 100,000, which cannot be carried, and implementations clamp it. - **Or set costs by hand** where the bandwidth is not the right measure, for example to encode latency or design intent. ## What a strong answer adds A strong candidate separates the three layers cleanly: the **protocol** carries an operator-chosen positive cost; the **implementation** offers a bandwidth formula with a default reference; the **operator** must make that reference consistent and large enough. They also know cost is static: OSPF never measures load, so a busy 10G link keeps its cost no matter how congested it gets.

  • If only some OSPF routers raise their reference bandwidth, does the network loop?
    No lasting loop. Each router advertises the cost of its own outgoing interfaces in its router-LSA, and every router runs SPF over the same link-state database, so all agree on the same costs. The damage is preference: unchanged routers still advertise 1 on fast links while updated ones advertise 10 or 100, so paths skew toward the unchanged routers and a flow's two directions can take different paths.
  • Is there an upper limit on how far an OSPF reference bandwidth can usefully be raised?
    Yes. A router-LSA carries interface cost in a 16-bit field, so 65,535 is the largest value. With a 100 Gb/s reference a 10 Mb/s link costs 10,000, which fits, but a 1 Mb/s link computes to 100,000 and cannot be expressed; implementations clamp it, and links that clamp to the same value become indistinguishable again. Choose a reference that keeps the slowest link inside the range.

saying these in an interview costs you the question

  • RFC 2328 defines OSPF cost as 100 Mb/s divided by interface bandwidth.
  • A 10 Gb/s link always gets a lower OSPF cost than a 1 Gb/s link.
  • Raising the reference bandwidth on one router fixes path choice for the whole area.
  • OSPF raises a link's cost automatically when the link gets busy.
  • Routers using different reference bandwidths form a permanent routing loop.
open as a page

When a link fails in an OSPF network, what steps make up the time until traffic takes a new path, and which usually dominates?

level: seniorimportance: must knowfreq 34%

basics

~20 s

OSPF convergence is detection, LSA origination, flooding, SPF and route install. Detection usually dominates: lost carrier is noticed at once, but a failure hidden behind a switch waits for the Dead interval, 40 s with RFC 2328's sample timers.

open as a page

Before maintenance on an OSPF router, how do you move transit traffic off it without dropping packets, and why not simply shut OSPF down?

level: seniorimportance: should knowfreq 20%

basics

~20 s

Cost the OSPF router out first: reoriginate its router-LSA with every transit link at the maximum cost, 0xffff (RFC 6987 stub router advertisement), let traffic reconverge around it, then work. A hard shutdown drops traffic until the area reconverges.

open as a page

How does OSPF graceful restart (RFC 3623) keep traffic flowing while a router's OSPF software restarts, and what ends it early?

level: seniorimportance: should knowfreq 15%

basics

~20 s

In OSPF graceful restart, the restarting router announces a grace period in link-local grace-LSAs and keeps forwarding on its preserved table; helper neighbours keep advertising it as fully adjacent until it resynchronises, the period expires or the topology changes.

open as a page

How would you design OSPF link costs for a network mixing 100 Gb/s core links, 10 Gb/s access links and slow WAN circuits?

level: principalimportance: nice to knowfreq 8%

basics

~10 s

Choose one deliberate OSPF cost scheme and apply it everywhere: a reference bandwidth above the fastest link that keeps the slowest inside the 16-bit cost field, or role-based costs that encode latency and intent.

open as a page