A JFR flame graph shows 70% of samples inside a small, three-line utility method called from dozens of hot call sites, and it's compiled at tier 4 early in the run. A team is about to spend a week rewriting it. What should they check before trusting that number, and why does this specific method — small, aggressively inlined, fully optimised — deserve extra suspicion rather than less?
They should check whether the profile's sampling mechanism itself is biased toward or against this specific kind of code before believing the number, and — counterintuitively — a small, heavily inlined, tier-4-compiled method is exactly the case where that bias is most likely to distort the result, not least. Some sampling profilers can only capture a thread's stack safely at a safepoint, a point in compiled code where the JVM knows precisely where every reference and frame boundary is, and HotSpot doesn't insert safepoint-poll opportunities uniformly through every instruction — tight, well-optimised, aggressively inlined code can run for a comparatively long stretch between poll points. A profiler forced to wait for the next safepoint over-samples whatever happens to be running at safepoint-friendly moments and under-samples code running in the gaps, which can make an innocent method look artificially hot. The team should double-check with an inclusive, call-tree view (is 70% really attributed to this three-line method, or is it really the caller it's inlined into?), and ideally cross-check with a second profiling mechanism, before rewriting anything based on one number.