In a gctrace line, what do `42->44->21 MB` and the `3%` field mean?
answer
- three heap sizes, one cycle
- start, end, and what survived
- the percentage is cumulative, not per cycle
- only one of the three tracks retention
basics
~20 sThey are the heap size when the cycle started, the heap size when it ended, and the live heap that survived marking. The percentage is the cumulative share of the program's CPU time spent collecting since start-up.
solid answer
~50 sThe triple `42->44->21 MB` reads as heap-at-start, heap-at-end, live-heap. The middle number is larger than the first because marking runs concurrently with the program, so the mutator keeps allocating while the cycle is in flight. Only the **third** number is a retention signal: it is what marking proved still reachable, so a live heap that climbs cycle after cycle is the thing that means a leak, while the first two mostly reflect where the collector was triggered. The `3%` is the fraction of the program's total CPU time spent in garbage collection **since the process started** — cumulative, not per cycle. That matters when you read it: a service that behaved for ten hours and then went pathological will still show a low percentage, so recent degradation shows up in the phase timings and the spacing of the cycles long before it moves that number.
code
text · 9 linesgc 47 @12.043s 3%: 0.11+5.8+0.09 ms clock, 0.9+1.4/4.2/0+0.7 ms cpu, 42->44->21 MB, 43 MB goal, 1 MB stacks, 0 MB globals, 8 P
# gc 47 cycle number since process start
# @12.043s seconds since process start
# 3% GC share of total CPU time, cumulative since start
# 0.11+5.8+0.09 ms STW sweep termination + concurrent mark and scan + STW mark termination
# 42->44->21 MB heap at cycle start -> heap at cycle end -> live heap after marking
# 43 MB goal heap size the collector aimed to finish under
# 8 P logical processors the runtime was usinggo deeper
Learn to point at the three heap sizes and say which is which, and to recognise that the percentage is about CPU time, not memory. Being able to spot the live heap in a line is the bar here.
Explain why the middle heap size exceeds the first — concurrent marking with the program still allocating — and separate the two stop-the-world clock figures from the concurrent one. That mechanical account is what this level is tested on.
Demonstrate that you read a series and not a line: rising live heap, cycles crowding closer, mark time growing. Call out the cumulative-percentage trap before an interviewer sets it for you.
Own the interpretation others will act on. Decide which of these figures becomes an alert, and insist that a cumulative average is never the number a pager fires on.
## The whole line A complete gctrace line has this shape: ``` gc 47 @12.043s 3%: 0.11+5.8+0.09 ms clock, 0.9+1.4/4.2/0+0.7 ms cpu, 42->44->21 MB, 43 MB goal, 1 MB stacks, 0 MB globals, 8 P ``` Each group answers a different question, and the two asked about here are the two people misread most often. ## `42->44->21 MB` — three heap sizes, not one The triple is **heap size at the start of the cycle**, **heap size at the end of the cycle**, and **live heap** — the bytes marking proved were still reachable. The first number tells you where the collector fired. It is a trigger point, not a measure of your program's appetite for memory: the collector starts a cycle when the heap reaches its target, so this number tracks the target far more than it tracks your data. The second number is almost always larger than the first, and that is expected rather than alarming. Go marks concurrently: the program keeps running and keeps allocating for the whole marking phase, so the heap grows underneath the collector. The gap between the first two numbers is roughly how much the program allocated while the cycle was in flight. A cycle that takes a long time to mark, on a program that allocates fast, shows a big gap. The third number is the one that matters for retention. It is the live heap: the memory that survived. If your program is genuinely accumulating memory it cannot release, this number climbs, cycle after cycle, and drags the trigger point and the collection frequency up behind it. If it is flat across hundreds of cycles, whatever else is going on, the reachable heap is not growing. A common mistake is to read the third number as "the memory this process uses". It is not. It is the live Go heap at one instant, measured immediately after marking. The process also holds goroutine stacks, runtime metadata, and heap spans the runtime has emptied but not handed back to the operating system, none of which appear in that number. ## `3%` — cumulative, not current The percentage is the share of the program's total available CPU time that has gone into garbage collection **since the process started**. It is an average over the whole run, not a reading for the cycle on that line and not a rolling window. The consequence is worth stating explicitly, because it catches people during an incident. A service that has been up for ten hours accumulated ten hours of denominator. If it began collecting furiously twenty minutes ago, those twenty minutes are diluted by the ten calm hours before them, and the percentage barely moves. Reading it as "we are fine, GC is only at 3%" during an incident is a real and common error. The way to get a recent figure is to difference two samples: note the percentage and the timestamp on one line, note them again some minutes later, and reason about the change. Failing that, the honest recent signals are elsewhere on the line: the phase timings getting longer, and the `@…s` timestamps of consecutive cycles crowding closer together. ## The other fields, briefly `gc 47` is the cycle counter and `@12.043s` the seconds since start-up; together they give you collection frequency for free — subtract consecutive timestamps. `0.11+5.8+0.09 ms clock` is wall-clock time for the cycle's three phases: **stop-the-world sweep termination**, **concurrent mark and scan**, and **stop-the-world mark termination**. The first and third are the pauses; the middle is concurrent work overlapping your program. So in that example the program was actually stopped for about 0.2 ms in total, even though the cycle spanned nearly six milliseconds. `0.9+1.4/4.2/0+0.7 ms cpu` is CPU time for the same work summed across processors. The middle three figures split the concurrent phase into assist time (collection performed inline with allocation), background collector time, and idle time. `43 MB goal` is the heap size the collector was aiming to finish under for that cycle. `1 MB stacks` and `0 MB globals` report the scannable goroutine-stack and global-variable memory. `8 P` is how many logical processors the runtime was using. ## Reading the series One line is a snapshot. The value is in a stretch of them, and there are three trends worth naming. A rising third number means real retention. Falling gaps between timestamps mean cycles are crowding — the program is allocating faster or the heap target is not moving up with it. A growing middle clock figure means marking itself is taking longer, which usually follows from more live pointers to trace.
- Why is the second heap size in the triple almost always larger than the first?Because marking is concurrent. The program is not stopped for the mark phase, so it keeps allocating while the collector traces, and the heap grows underneath it. The gap between the first two numbers is approximately what the program allocated during the cycle. A wide gap means either a long mark phase or a fast allocator, and it is normal rather than a fault.
- A service degraded twenty minutes ago but the gctrace percentage still reads 3%. Why?That field is cumulative over the whole process lifetime, so twenty bad minutes are averaged against however many good hours preceded them. To get a recent figure, difference the percentage and the timestamp between two lines some minutes apart. Meanwhile the timely signals are the growing mark-phase clock time and consecutive cycle timestamps crowding closer together.
- In `0.11+5.8+0.09 ms clock`, how much time was the program actually stopped?About 0.2 ms — the first and third figures only. Those are the two stop-the-world phases, sweep termination and mark termination. The 5.8 ms in the middle is concurrent mark and scan, which runs alongside the program rather than pausing it, so quoting the whole 6 ms as pause time overstates the impact by roughly thirty times.
saying these in an interview costs you the question
- Reads the live-heap number as the process's total memory use
- Takes the percentage as the GC cost of the last cycle
- Calls the whole clock triple stop-the-world pause time
- Treats a rising first number alone as proof of a leak
- Assumes the heap cannot grow while a cycle is running