~/tech-with-ugur

Three Valleys, One Gradient Step

2026-09-29 maths

Run the companion lab

Picture a ball on a landscape with three valleys. A rule that always moves downhill will find a valley, but which one depends on where the ball starts. Gradient descent has the same limitation when its cost function has several local minima. In one dimension, we can see every move rather than hiding it inside a large model.

The lab runs the same 100-update loop from x = -2.6, 0.5, and 2.6. It writes a two-panel plot and complete CSV and Markdown tables. The function, derivative, starting points, and learning rate are fixed, so the paths can be reproduced with docker compose up --build from the lab directory. No account or external data is needed.

A cost with three valleys

The cost is f(x) = x²(x² − 4)² / 16 + x / 10. The squared factors give the curve several low regions; the x / 10 term tilts them so they do not all have the same depth. Expanding gives x⁶/16 − x⁴/2 + x² + x/10. Differentiating term by term gives f′(x) = 3x⁵/8 − 2x³ + 2x + 1/10, or, in factored form, x(x² − 4)(3x² − 4)/8 + 1/10. The lab evaluates that known slope directly:

labs/lab-gradient-descent-1d/src/app/experiment.py:

def derivative(x: float) -> float:
    """Evaluate the analytical derivative at a finite point."""
    if not _is_finite_real(x):
        raise ExperimentError(f"x must be finite, got {x}")
    value = x * (x * x - 4) * (3 * x * x - 4) / 8 + 0.1
    if not math.isfinite(value):
        raise ExperimentError(f"derivative is non-finite at x={x}")
    return value

The derivative tells us the direction of the local slope. At a point with a positive slope, subtracting it moves x left; a negative slope moves x right. With learning rate 0.04, each update is x_next = x − 0.04 × f′(x). The rate controls the size of each move, while the derivative supplies its direction and scale.

One row tells you the next

The loop records the current position, cost, and slope before each update. That matters when reading the table: the slope in row k determines the position in row k + 1.

labs/lab-gradient-descent-1d/src/app/experiment.py:

        for iteration in range(steps + 1):
            row = _record_state(run, iteration, x)
            rows.append(row)
            if iteration < steps:
                x -= learning_rate * row.slope

Here are the first four rows of the left run, copied from the lab’s generated output/iterations.md:

runiterationxcostslope
left0-2.62.9584360000000016-14.503160000000005
left1-2.0198736-0.200359915125112-0.06620070631879924
left2-2.017225571747248-0.20050472154481033-0.0432123158076641
left3-2.0154970791149416-0.200566551326031-0.02834810285373976

At iteration 0, the slope is about -14.50316. Substituting it into the update gives -2.6 − 0.04 × (-14.50316) = -2.0198736, exactly the next row’s position. The next slope is much smaller, so the following move is only about 0.002648. The left path has already reached the neighborhood of its valley after one large step; the remaining updates refine its position.

Each run records iterations 0 through 100, so the two tables each contain 303 data rows. Start with the readable output/iterations.md; use output/iterations.csv when you want the full values for your own checks.

Three starts, three endings

Two-panel gradient descent plot. On the left, blue, orange, and green paths from -2.6, 0.5, and 2.6 settle in three different valleys of the same cost curve. On the right, their costs fall over 100 iterations and level off near -0.201, -0.003, and 0.199.

The left panel draws every recorded position on the cost curve. Many late points overlap where a path settles. The right panel uses iteration on the horizontal axis, so those later steps remain visible even when the position barely changes.

After 100 updates, the left run reaches approximately x = -2.012164 with cost -0.200614. The center run reaches x = -0.049972 with cost -0.002503; the right run reaches x = 1.987131 with cost 0.199363. The left start found the lowest observed result of the three. Starting in the center or right basin did not carry the same update rule across the hill to the left valley.

These endings are not merely a plotting effect. The lab records all three paths through one call pattern:

labs/lab-gradient-descent-1d/src/app/commands/run.py:

        rows = tuple(
            row
            for name, start in (("left", -2.6), ("center", 0.5), ("right", 2.6))
            for row in gradient_descent(start, 0.04, 100, run=name, log=log)
        )

Only the start changes. Each run uses the same derivative, learning rate, and number of updates.

What the comparison proves

Trying several starts is useful: it finds a better result here than either the center or right run alone. It does not, by itself, prove that no lower valley exists elsewhere. Finite gradient descent runs also approach a minimum numerically; they do not reach an exact stationary point by declaration.

For this particular polynomial, there is a separate global-minimum argument. Its derivative has five real stationary points: three local minima near -2.012, -0.050, and 1.987, separated by two local maxima. A degree-five derivative can have no more than five real roots, and the positive x⁶ leading term sends the cost to positive infinity as |x| grows. Comparing the costs at those three minima identifies the left one as the global minimum of this function. That mathematical check supports the example; multi-start gradient descent has no general global-optimum guarantee for nonconvex costs.

One dimension makes the update easy to inspect: one number, one analytic slope, one recorded move at a time. A later lab can extend the same idea to multiple parameters and a numerically estimated derivative, where the path is harder to picture but the question of where a start leads remains.

Run the complete lab to inspect every table row and reproduce the plot.