What a Lollipop Taught Me About Geometry
A friend asked whether you could lick a round lollipop into the shape of the Earth. The optimal strategy turns out to be one line long, and getting there cost me a week of minimising the wrong thing.
English 中文updated 13 Sept 2026 18 min read #differential geometry #holonomy #geometric control
Can you lick a perfectly spherical lollipop into the shape of the Earth?
A friend asked me this, and my first answer was wrong in a way I did not notice for a while. Total the volume you need gone, divide by the volume one lick takes, done.
Ignore mountains and oceans. At planetary scale the Earth is an oblate spheroid: a little wider around the equator than from pole to pole. Shrink that same proportion onto a round sweet and three questions fall out.
- How much candy has to go?
- Where does it come off?
- If a lick leaves a long thin mark rather than a dot, which way should the mark point?
The first is arithmetic and I will do it in a paragraph. The second is harder. The third one is the reason this page exists, and I did not see it coming.
The model here is a fixed sphere, additive removal, and no saliva. Real candy dissolves, the tongue is warm and wet, and the shape changes while you work on it. Everything after this sentence is about a sphere that does not care.1
How much, and then where
The flattening of a spheroid is the fractional gap between its equatorial and polar radii,
WGS84, the model your phone’s GPS uses, puts the inverse flattening at 298.257223563. So . On a lollipop of radius mm that makes the polar dimple 67 µm deep — about the thickness of a sheet of paper. Keep the equator where it is, pull the poles in by that much, and the volume you owe is
which is 0.34% of a 33.5 cm³ sweet. That was my first answer, and it is a correct answer to a question nobody asked. A tin of paint knows its own volume and has no opinion about the painting.
Remove that much evenly and you get a smaller sphere. Remove more near the equator and you have made an American football. What the problem actually wants is a field: a rule saying how deep to go at every point.
Let be colatitude, measured down from the north pole. Expanding the target spheroid to first order in ,
Check the ends. At a pole and you take the full 67 µm. At the equator it is zero and you take nothing. Geometry has answered the question the volume calculation could not even hear.2
Two notes
Spherical harmonics are the notes of a sphere. A pattern painted on a ball decomposes into a fundamental and its overtones the way a struck string does, and the low harmonics are the broad patterns while the high ones are the fine detail.
Our target is a two-note chord. is the polynomial you want:
A constant, which is degree zero, plus one broad pole-versus-equator pattern, which is degree two. Nothing else survives the linearisation.
Two notes. That should be the easy case, and for a while I thought it was going to be, because the instrument is the problem: a single lick is a small patch, and a small patch is a chord of dozens of harmonics at once. You cannot play degree two on its own. You can only play thick clusters and hope the unwanted partials cancel.
Except they can’t cancel, because a lick removes candy and cannot put it back. Every amplitude is non-negative. That is where the music analogy dies, and it dies at the point that makes this problem interesting: I get to choose where to strike, every strike is positive, and there is no such thing as playing a note at negative volume to silence the one before it.
The week I spent minimising the wrong thing
So I set up the obvious objective. Pick a menu of possible licks, give each one a non-negative strength , and minimise the squared distance from the target field:
Twelve contact points spread over the sphere, four orientations each, 48 candidates, non-negative least squares. Relative error 0.4365. I tried more orientations. I tried moving the centres. Nothing got below about 0.4, and I spent a week convinced the obstruction was deep.
The obstruction was that I was minimising the wrong thing.
Go back and look at what a constant does. A uniform layer of removal takes the same depth off everywhere. It shaves the lollipop down to a smaller lollipop, and a smaller lollipop is the same shape. The degree-zero harmonic is not part of the design problem at all. It is a free parameter that costs candy and buys nothing, and I had been faithfully fitting it, at full weight, for a week.
Only degree two matters. Of my 0.4365, a large chunk was the essay-grade error of insisting that equal on the nose, when what was wanted was equal to plus anything constant.
The whole answer, in one line
Drop the constant from the objective and the problem stops being an optimisation. It becomes a division.
Here is the move. The target depends only on , so look for a removal recipe that also depends only on : a density over the sphere, uniform in longitude, licking in rings. Smearing a footprint evenly around a circle of latitude and then adding up the rings is a spherical convolution, and convolution on a sphere is multiplication note by note. If the footprint has degree- transfer
then a density with note produces a field with note . Every note is independent. To get degree two right, divide by .
Write . Matching degree two needs . Degree zero is free, so make it as small as non-negativity allows: bottoms out at , so , and
The optimal lick density has the same shape as the thing you are carving. Lick hardest at the poles, taper as , stop entirely at the equator, and spread each ring evenly around its circle. The resulting removal field is
which is the target plus a constant. Shape error zero. I checked it numerically and got , which is what “zero” looks like when you compute it in double precision.
For the footprint I have been using, .3 So the recipe removes times the bare volume difference: 133 mm³ instead of 112. The extra 18% is not waste in any interesting sense. It goes entirely into that uniform term, and on a 20 mm lollipop the uniform term is 4.0 µm.
Four microns. You are carving a 67 µm feature and the price of doing it exactly right is a lollipop four microns smaller than the theoretical minimum. Nobody is going to notice four microns. That is the whole answer to my friend’s question, and it took a week of not looking at it.
Three hundred licks
A density is not a strategy. You cannot lick continuously; you lick a finite number of times. So put rings at the Gauss–Legendre colatitudes, give ring the mass its share of demands, and split that mass evenly among licks spaced around the circle, with scaled to the ring’s circumference so the licks overlap by about the same amount everywhere.
| Rings | Licks | Shape error |
|---|---|---|
| 5 | 126 | |
| 7 | 192 | |
| 9 | 258 | |
| 11 | 316 | |
| 13 | 384 |
Eleven rings and 316 licks gets the shape to five parts in a million. Solve the ring masses again by non-negative least squares, throwing away the formula, and you recover the same numbers to eight digits. That is the check I actually trust.
Three hundred licks is also, and I did not plan this, roughly how many licks it takes to finish a lollipop.
What the optimal strategy actually tells you to do
Lick in rings. Each ring of latitude gets a total amount of removal proportional to cos² of its colatitude, split evenly around the circle. Drag the slider to add rings; watch the solid curve settle onto the dashed one.
11 rings, 316 licks, shape error 5.0e-6. The solid curve sits above the dashed one by a constant, and a constant is the one thing that does not matter: it shaves the lollipop down evenly without touching its shape. That offset is why the recipe removes 1.180× the bare volume difference. The equatorial ring gets nothing, so it is not drawn.
Drag it up to eleven rings and watch the solid curve settle onto the dashed one, displaced upward by a constant. The constant is the whole lesson.
The state of a lick
Everything so far assumed the footprint is round. Time to spend the rest of the essay on the assumption I quietly made and my friend’s question did not grant me.
Begin with a circular contact patch. To place it you need a centre and a size; rotating it about its own centre does nothing at all. Now stretch it into an ellipse. Two patches with the same centre and the same two widths, one lying north–south and the other east–west, remove different candy. A point on the sphere has stopped being enough to say where you licked.
At the contact point, lay down two little perpendicular arrows, one along the long axis and one along the short. Both lie flat on the candy; nothing pokes through. They are tangent vectors. Together with the outward normal they make a full three-dimensional orientation, which is to say a rotation matrix:
The last column is where you are and the first two are which way the lick points. The oriented orthonormal frame bundle means nothing more alarming than “every point of the sphere, together with every pair of perpendicular unit arrows you could attach there,” and if you erase the arrows and keep only the position, many frames collapse onto one point:
I am carrying more information than the ellipse has. An ellipse turned through 180° is the same ellipse, so the frame distinguishes states the footprint cannot. That is sloppy, and it is also the only reason transport is easy to write down, so I keep it.
For the footprint itself: strongest at the centre, fading outward, stretched into an ellipse. Unroll a small neighbourhood onto the tangent plane with the logarithm map, work in polar tangent coordinates , and write
You can read the parameters off without doing the integral. sets the average narrowness; says how unequal the two widths are. At the mark is round and orientation is irrelevant. As grows it gets longer and thinner and cares more about which way it is turned. Throughout, and , so the footprint’s width is rad, which is 5 mm of arc on a 20 mm lollipop. That is about a tongue.
The factor is a smooth cutoff that keeps the kernel inside a cap:
zero outside. It is there because the logarithm map has no opinion about direction at the antipode, and rather than argue with the antipode, mathematics just went ahead and truncated. It is a fudge. I have never liked it. The normalisation makes each footprint integrate to one over the sphere, using the spherical area element and not the flat one.
That is everything needed to write down what happens when a lick gets turned.
Parallel transport
Slide an arrow along the lollipop. Keep it flat against the surface and never deliberately twist it left or right. On a unit sphere sitting in ordinary space the rule is one line:
All the change is along the normal and none of it is in the tangent plane, which is the formal content of “no twisting.” Differentiate and you find tangency is preserved; differentiate and you find length is too.
Now walk a full circle of latitude. You come back to the same point with the arrow pointing somewhere else. At colatitude 45° it is off by 105.4416°, or 1.840302369021 radians if you want it that way, and the twelve digits are not showing off. They are the cheapest known way to catch an integrator bug, which is the only reason anyone ever prints twelve digits. Gauss–Bonnet says the angle is the enclosed cap area,
and the honest way to say what went wrong is that “without twisting” is a condition on each tiny step and not on the trip. Every step was locally straight. Curvature is the fact that locally straight choices do not fit together globally.
Nothing slipped. The sphere is curved, and that is the whole of it.
How much damage a turn does
An angle is not yet a consequence. Rotate a round lick by any amount you like and it removes exactly the same candy, which means the holonomy angle I worked so hard for can be, for the wrong footprint, completely invisible. So: how different are two licks?
Compare the footprint you left with the footprint you would leave now, same centre, and integrate the squared difference:
The overlap is easier, and since rotation preserves the norm, . Multiplying the two bell curves puts two cosines in one exponent, which collapse:
The angular average of is the modified Bessel function , so a two-dimensional surface integral turns into one integral over distance from the contact centre:
That is the formula. Read it for a moment before I tell you what is in it.
is even and increases with the size of its argument. So vanishes at and , peaks at , and grows with . Which means the worst route is the one with the most holonomy: walk a big loop, come back badly turned, leave a badly wrong mark.
Except at 60°.
At colatitude 60° the enclosed cap has area exactly , so the frame comes back rotated by a full half-turn. Maximum possible insult. And the footprint is identical, because turning an ellipse through 180° maps both of its axes onto themselves reversed and an ellipse cannot tell. I spent an afternoon convinced the quadrature was broken. It was not broken. The geometry moved as far as it can move and nothing measurable happened at all.
I do not know whether that coincidence generalises, and I could not find it stated anywhere.
Carry a lick around the sphere and it comes back turned
Walk the contact once around a parallel without ever twisting your hand. It returns to the same spot pointing somewhere else. The curve on the right is the exact mismatch between the footprint you left and the footprint you would now leave.
A round footprint (elongation 0) is indifferent to all of this. Everything on this panel is the price of having a long axis.
The mismatch in that panel is the exact formula above, evaluated in your browser as you drag. Put the route at 60° and watch it go to zero while the frame rotation reads 180.00°.
One more thing the formula gives away. Expand the Bessels in small ; the constants cancel when you subtract, and the leading term is quadratic:
So the mismatch is linear in the elongation, while its square is quadratic, and keeping those two straight is a small piece of bookkeeping that will otherwise cost you an afternoon. Writing and taking the 45° loop, the expansion predicts at leading order:
| Relative mismatch | |
|---|---|
| 0.01 | 0.008618 |
| 0.02 | 0.017236 |
| 0.10 | 0.086272 |
| 0.40 | 0.350946 |
| 0.60 | 0.539096 |
Doubling the smallest contrast doubles the mismatch. The measured log slope is 0.999973. I wanted 1. The gap is quadrature error and it is boring, which is exactly what you want from that gap.
What a long axis is actually worth
Here is where I expected orientation control to pay for itself, and it does the opposite.
Put the elliptical footprint back into the ring recipe. The closed form assumed a round kernel, because a round kernel makes the ring construction a genuine convolution. A tilted ellipse does not: the effective degree-two transfer now depends on the colatitude of the ring and on the tilt of the long axis, as . At 45° colatitude, a long axis lying along the meridian delivers 0.2234 of degree-two response per lick against 0.1871 for one lying along the parallel. Meridian wins by 19%.
Then you fit it and the meridian rule loses badly.
| Footprint at each site | Licks | Shape error |
|---|---|---|
| one tilt, along the meridian | 200 | |
| four tilts, 0/45/90/135° | 1264 | |
| eight tilts | 2528 | |
| round | 316 |
A fixed tilt leaves azimuthal order in every footprint. The target has no in it, eleven ring strengths cannot cancel it, and going to 21 rings only gets you to . It is a floor, not a convergence rate.
Four tilts in equal measure kill orders 2, 4 and 6 exactly (the four phases sum to zero unless the order is a multiple of eight), and the first survivor is , sitting at of the footprint. Which is, to within a factor of two, the error in the table. Eight tilts push the survivor out to and the error falls by another factor of twenty-five, at which point ring spacing is the limit again and anisotropy has stopped mattering.
The best thing anisotropy can do here is cancel itself
Same eleven rings, same target, same non-negative fit. The only difference is how many tilts the tongue is allowed to use at each site. More tilts is worse value per lick and better value per answer.
Tilts at 0°, 45°, 90° and 135° in equal measure annihilate azimuthal orders 2, 4 and 6 exactly. The first survivor is m = 8, and that residue is precisely what sets the error here.
Averaging over four tilts leaves a footprint that is rotationally symmetric but radially fatter, with an factor smeared into its profile, so drops from 0.847721 to 0.820925 and you pay 1.2181 times the volume instead of 1.1796. That is the real cost of having a long axis: not a worse shape, just more candy.
So for a target this smooth and this symmetric, directionality is a liability, and the best thing an orientable tongue can do is arrange to cancel its own orientation. That is not what I expected to find, and I want to be careful about how far it generalises: a zonal degree-two target is about the friendliest thing you could ask for. Ask for something with genuine directional structure and the accounting inverts.
Where the bill arrives
Which brings back the arrow.
Suppose you do commit to the meridian rule, because you have one tongue and it has one shape. You now have to hold the long axis on the meridian all the way around a ring, and that is not free. The parallel at colatitude has geodesic curvature and arc length per radian of longitude, so holding a fixed bearing costs you
of deliberate twist. Over a full lap that is , which at 45° is 254.56°. And what you decline to supply, curvature supplies for you: the remaining is exactly the holonomy, 105.44°, the same number from three sections ago.
Foucault measured that number in 1851, in a church, with a 67-metre wire and a 28-kilogram bob. A Foucault pendulum at latitude precesses relative to the ground by per day, and , so the precession he watched is exactly the twist you have to supply. The Earth turns a full circle; parallel transport keeps of it; Foucault’s pendulum shows you the change.
The holonomy formula is not a curiosity that happens to live near this problem. It is the correction term in the control law. Lick around a ring with a directional tongue and you are either paying per radian or you are accepting of drift, and there is no third option, because the sphere is curved.
Anyway, it’s a lollipop.
What is actually proved
Three checks. The second one is the only one I would defend in a seminar.
The first is routine: integrate the transport equation with fourth-order Runge–Kutta, halve the step, watch the error fall like the fourth power. A 20,000-step run lands within rad of the analytic angle. The second computes the mismatch twice, once through the one-dimensional Bessel formula and once by brute-force integration of the squared difference of the actual kernels over a radial and angular grid, with two independent normalisations. Worst disagreement in the table: . That is double precision complaining, not a disagreement. The third checks the non-negative least-squares optimality conditions and re-evaluates the fitted strengths on a finer sphere grid, so the reported errors are not a pixel-resolution illusion.
What none of that establishes is a route. The recipe is a spatial prescription: lick here, this hard, this many times. It says nothing about what order to do it in, and with a directional tongue the order matters, because transport changes the arrows you carry and therefore changes every footprint after the first. A tool that resets its own orientation at each contact makes the history dependence go away. A tongue does not.
Prior work
Almost none of the geometry here is new, and it is worth being specific about which parts. Directional wavelets on the sphere go back decades, and Antoine and collaborators give an early implementation and analysis of orientation-sensitive functions on . Cohen and collaborators make local frames explicit in gauge-equivariant convolution and are clear that transported filters are path dependent, which is the machine-learning version of the thing this essay is about. Schonsheck, Dong, and Lai build parallel-transport convolution on manifolds directly. The fact that makes the whole answer collapse into a division is older than any of it: the Funk–Hecke theorem, 1916, which says a zonal kernel acts on each spherical harmonic by multiplication, with the multiplier given by exactly the integral above. Meanwhile, in a literature that has never heard of any of this, Tam and Cheng show experimentally that polishing paths and elliptic contacts change material removal, which suggests the practitioners got to the important part first, as they usually do. Do Carmo’s chapter on the intrinsic geometry of surfaces does parallel transport in general if you want it without the sphere’s special structure. What I have not found anywhere is the particular bridge: the density as the exact non-negative optimum once you stop fitting the constant, the explicit one-integral orientation response for this kernel, the 180° exception, and the fact that four tilts are rotationally symmetric through order seven and that this is why they work.
There is a technical paper, and the
code
reproduces every number above; code/optimal_strategy.py is the one that does
the strategy.
So, can you?
Yes. Lick in rings, hardest at the poles, tapering as , nothing at the equator, about three hundred times, and keep your tongue round or keep it turning. You will overshoot the minimum removal by 18% and finish four microns small, and the shape will be right to five parts in a million.
What I got wrong at the start was not the arithmetic. It was thinking the answer lived in the amount. The amount is one number and it is the least interesting number in the problem. Everything that mattered was in the second question, and the whole third question turned out to be about a bill that curvature sends you for having a preferred direction, which is a thing I did not know a lollipop could be about.
Footnotes
-
There is one caveat I cannot fold into that sentence, because it is load-bearing: removal is assumed additive, so two overlapping licks remove the sum of what each would remove alone. Real dissolution saturates. Everything in this essay is a linear-response calculation about the first instant. ↩
-
At second order in a term appears, with coefficient , which brings degree four into the target. The method in the next section handles it without modification: add a note to the density and divide by instead of . The correction is 0.34 µm on a 20 mm lollipop. I mention it only so you know the linearisation is a choice and not a limitation. ↩
-
depends on how wide the footprint is, and only on that. A narrower tongue is closer to a delta function, so it smooths less and rises towards 1: at it is 0.955354 and the overshoot drops to 4.7%. The cost of a wide tongue is entirely a cost in candy, which is a pleasant thing for a cost to be. ↩