Questions to ask when benchmarking clocks
You have two or more clocks and you want to know which is better. The awkward part is that a comparison is itself a measurement, and it can be wrong in all the ways the clocks can be wrong — quietly.
1. Is your reference clock better than the clock you are measuring?
You cannot measure a clock with a worse clock. Your reference clock contributes its own error to every measurement you make. Is it so much better that you can neglect that error? If not, your benchmark results have to be qualified by the error of the reference.
An analogy: you can use a digital wall clock that displays seconds to benchmark one that shows only hours and minutes. But it is very hard to say anything conclusive when comparing two wall clocks that both show only hours and minutes. It is an imperfect analogy, because it may suggest that resolution is the only source of error — but it might still help you explain the challenge to your boss.
2. Are you measuring trueness or precision?
They want different setups and answer different questions. Precision you can measure against any stable reference, even a wrong one. Trueness needs a reference you believe, which in this field means a traceable one — and, because UTC arrives later, possibly a wait.
3. Is your counter’s resolution finer than the differences you care about?
If it is not, every clock will look identical and excellent. This is the counterfeit precision trap and it is especially easy to walk into here, because the instrument reports a confident number either way.
4. Over what averaging interval — and are you comparing the same one?
Stability is a curve. Two clocks can trade places depending on where you read it: the one that wins at 1 s may lose at 10,000 s. Comparing single numbers taken at different intervals is not a comparison at all.
5. Common reference, or pairwise?
Measuring everything against one reference is simple and puts the reference’s error into every result equally. Measuring pairwise lets you separate the units from the reference — the three-cornered hat — but needs three clocks and assumes their errors are uncorrelated, which is exactly what a shared environment breaks.
Put two clocks on a two-channel time interval counter and it is natural to think you are comparing two things. You are comparing three. The counter digitizes both channels against its own timebase — its reference clock — and that oscillator is in every reading it gives you.
It is easy to miss because the counter reports a confident number either way, and because the timebase is usually inside the instrument where nobody looks at it.
So run a three-cornered hat on anything measured this way. You have the three clocks whether you wanted them or not, and the analysis will tell you which of the three you can actually trust — including the one you were treating as furniture.
And watch where the counter’s reference comes from. If you feed its external reference input from one of the clocks under test, that clock is now in the measurement twice, and you have broken the uncorrelated assumption with your own patch cable. The three-cornered hat will usually announce it, by returning a negative variance for something.
6. Have you separated the device under test from the measuring system?
Cables, connectors, the counter’s own trigger noise and any distribution amplifier are all in the loop, and none of them announce themselves.
The standard test for a two-channel time interval counter is a zero-baseline measurement, and it takes two minutes: split one 1 PPS with a tee connector and run it to both channels on carefully matched lengths of coax. Both channels are now looking at the same edge at the same instant, so the true interval is exactly zero. Everything the counter reports is the counter.
You get two numbers out of it, and they are different things:
- The spread of the readings is your counter’s single-shot resolution — the floor beneath which none of your results mean anything, no matter how many digits are displayed.
- The mean is the channel-to-channel offset, a systematic you can simply subtract from every subsequent measurement.
Matched cable lengths matter for the second one: any length difference lands directly in that mean and is indistinguishable from the instrument’s own bias.
7. Are cable delays calibrated out?
The same arithmetic that makes a feedline matter applies to every cable on the bench. A meter of unaccounted coax is several nanoseconds of systematic error assigned to whichever clock happened to be on the long side.
8. Is anything else different between the units?
Temperature, supply, load, position in the rack, time of day. A systematic difference in conditions produces a systematic difference in results, and it will look exactly like one clock being better. Swap the units between positions and re-run — if the ranking follows the position rather than the unit, you have learned something more useful than the ranking.
9. How long did you run?
Long enough to tell a bias from a drift, which is a question about the shortest timescale you care about and the longest you can afford. Short runs systematically flatter clocks whose errors are slow.
If the answers start raising questions about why a clock behaves the way it does rather than whether it does, you have crossed into designing one.