ELAI S.r.l.

Guiding an image generator: why more strength does not mean more quality

Classifier-free guidance changes the generation process. A worked Gaussian example separates local direction, final distribution and loss of variety.

Guiding an image generator: why more strength does not mean more quality

A precise request can produce overly similar results

A design team wants varied images of an object, all matching a description. It raises the generator’s guidance parameter for stronger adherence. Yet adherence, variety and visual quality differ: restricting possibilities can remove useful alternatives, and pushing a direction too strongly can overshoot the desired region. How can we understand the change without treating guidance as a knob labeled quality?

Abstract. Explain classifier-free guidance, CFG, as a combination of model-estimated directions. Then construct a one-variable example with explicit calculations: the distribution suggested at one instant differs from the distribution produced by the full trajectory. At γ=3 our static variance is 0.4, versus transported variance about 0.06265. This does not measure image beauty; it establishes a mathematical distinction needed to interpret the control. We connect it to recent research without calling our calculations visual-model benchmarks.

From noise to a direction: what a score means

A diffusion model learns to reconstruct progressively noise-corrupted data. No particular network is needed here: at noise level t, imagine density p_t(x) describing plausible values. Its score is the derivative of log density with respect to x, or a gradient vector in multiple dimensions. It is not an image-quality rating; it indicates local increase of log density.

score(x) = d log p(x) / dx p(x) = N(μ,v) ⇒ score(x) = −(x−μ)/v

N(μ,v) denotes a normal distribution with mean μ and variance v. Right of the mean the score is negative; left it is positive. Smaller variance gives a stronger direction at equal distance. Differentiate −(x−μ)²/(2v) in log density; the x-independent normalizer vanishes. This elementary case lets us separate a local direction from the distribution of points transported along it.

Guidance compares two descriptions of the same point

Let s_c be the direction conditioned on request c and s_u the unconditioned reference. Combining s_u+γ(s_c−s_u) amplifies the condition-induced difference. γ=0 uses the reference, γ=1 the conditional direction, and γ>1 extrapolates beyond it. Ordinary weighted averages have nonnegative weights; here reference weight 1−γ becomes negative. Guidance therefore involves extrapolation, not simply selecting existing samples.

s_γ = s_u + γ(s_c−s_u) = γ s_c + (1−γ)s_u

For s_u=−1 and s_c=0.5 in one coordinate, γ=3 gives −1+3·1.5=3.5. Quality was not tripled; a direction field changed. Parameter conventions matter: in (1+w)s_c−w s_u, our γ equals 1+w. Comparing guidance 3 across implementations without checking formulas can compare different operations.

A mathematical snapshot at fixed noise

For exact scores, the gradient combination equals the log-gradient of p_c(x)^γ p_u(x)^(1−γ). A density requires division by a finite integral Z. This identity holds at a fixed noise level; it does not establish that traversing all levels generates that final density. The local map and the complete journey are different claims.

r_γ(x) = p_c(x)^γ p_u(x)^(1−γ) / Z p_c = N(1,1), p_u = N(0,4) a_γ = γ + (1−γ)/4 μ_static = γ/a_γ, v_static = 1/a_γ

Choose synthetic Gaussians with dimensionless x: conditional mean 1, variance 1; reference mean 0, variance 4. Combining quadratic log terms gives coefficient a_γ on x² and linear term γx. Completing the square yields mean γ/a_γ and variance 1/a_γ. At γ=3 these are 1.2 and 0.4: even the static density narrows and shifts past 1. This is dispersion on an abstract axis, not semantic diversity.

Evolve the points, not just the formula

Add Gaussian noise of variance t: densities become N(1,1+t) and N(0,4+t). Here t is noise variance, not computation seconds. The combined score is s_γ(x,t)=−a(t)x+b(t), with a(t)=γ/(1+t)+(1−γ)/(4+t), b(t)=γ/(1+t). Use the probability-flow equation dx/dt=−s_γ/2, traversed from T=100 down to zero with negative time increments. Changing sign without reversing traversal describes another process.

dx/dt = (a(t)x − b(t))/2 dμ/dt = (a(t)μ − b(t))/2 dv/dt = a(t)v

The first equation moves each point. An affine x motion preserves Gaussianity, so tracking mean μ and variance v suffices. Mean follows the center’s motion. Deviations evolve with coefficient a/2; squaring introduces a factor two, yielding dv/dt=av. These equations describe transported density. We did not impose v=1/a(t), which belongs to the static snapshot instead.

For a controlled comparison initialize with the static density at T: μ(T)=b(T)/a(T), v(T)=1/a(T). This declared choice is not exactly an image generator’s common prior. At γ=3 it is approximately N(2.83636,95.49091). Different initializations by γ make the check explicit: even starting on the correct static snapshot does not force transport to follow every later snapshot.

The result can be checked without drawing samples

Integrating d log v/dt=a(t) from T to zero gives exact final variance. Ratio 1/(1+T) comes from conditional-variance logarithms, 4/(4+T) from reference variance. Exponents are guidance weights. All factors are positive here, avoiding power or normalization ambiguity.

v(0) = v(T) [1/(1+T)]^γ [4/(4+T)]^(1−γ)
γStatic meanStatic varianceODE meanODE variance
00404
11111
31.20.41.4972860.0626534

ODE means ordinary differential equation. We integrated mean and variance with fourth-order Runge–Kutta, step −0.025, and repeated at −0.05. At γ=3 the maximum difference is below 6.4×10⁻⁸; numerical variance differs from the analytic result by less than 4.1×10⁻⁹. γ=0 and γ=1 recover expected distributions, checking signs and initialization. No random sampling means no seed is needed. Python code and results are attached.

Gaussian densities at t=0, γ=3: desired conditional, normalized static product and distribution transported by the specified ODE. x is dimensionless; each curve integrates to one. Greater height indicates concentration, not greater quality.
Gaussian densities at t=0, γ=3: desired conditional, normalized static product and distribution transported by the specified ODE. x is dimensionless; each curve integrates to one. Greater height indicates concentration, not greater quality.

Transported variance 0.06265 is far below 0.4 and its mean overshoots further. Matching instantaneous scores alone therefore does not establish the final distribution. This proves neither universal collapse nor wrong images nor that γ=3 is always bad. We used Gaussians, exact scores and a particular continuous process. The useful distinction is that local field, sampling dynamics and output distribution require joint analysis.

What the papers add, and what we checked

Ho and Salimans, arXiv version dated July 26, 2022, combine conditional and unconditional estimates, randomly dropping conditions during training. They evaluate 64- and 128-pixel ImageNet with 50,000 samples per guidance setting. Their metric comparison is not a universal quality ranking.

Jiang and Ma’s August 6, 2026 preprint arXiv:2607.19725v2 studies trajectory corrections and proposes DG-CFG. DDIM evaluations use SD1.5, SD2.1 and SDXL: 1,000 COCO prompts, with diversity assessed on 100 prompts and 16 images each. These author results were not reproduced here.

Our verification covers only the stated Gaussian equations. We trained no network, compared no checkpoints and measured no prompt adherence. A recent preprint is research, not certification: exact-score theory does not remove learning errors, discretization or sampler differences. Automated metrics may reward one property while missing another. A business creative process may value useful alternatives more than one index; a constrained visualization may prioritize differently. Define the problem before choosing the parameter.

Designing a meaningful experimental comparison

For a real generator, hold model, version, prompt, resolution, sampler and network-evaluation budget fixed. Compare guidance strengths and separately a noise-dependent schedule. Reusing seeds across configurations supports paired comparison, but many seeds and prompts are needed: one successful image does not characterize a distribution. This protocol is proposed, not executed. Record failures, semantically distinct alternatives and constraint compliance alongside automated metrics.

Cost also needs care. Classifier-free does not mean no additional work: when both conditional and reference predictions are needed, both must be produced. Implementations may batch or reuse computation; step count alone does not describe time, memory or evaluation count. A negative-prompt reference is not automatically unconditional either: it changes the extrapolation reference. Such details can change an experiment while the displayed guidance number stays identical.

Conclusion: choose a tradeoff, not a maximum knob setting

Increasing guidance changes sample-building directions and can alter both center and dispersion. Our example showed something further: a valid instantaneous density formula may fail to describe the final distribution. Informed control requires specifying what to preserve, which dynamics run and how outputs are evaluated collectively. Quality does not automatically follow greater strength; it must be defined and checked for the task.

Bibliography and verifiable material

Jonathan Ho, Tim Salimans — Classifier-Free Diffusion Guidance, arXiv:2207.12598v1, 26 July 2022; sections 3–4.

Enze Jiang, Zheng Ma — Analytic Distribution of Classifier-Free Guidance for Schedule Design, arXiv:2607.19725v2, 6 August 2026, preprint; sections 4–6.

The short program calculates exact variance. In the archive, experiment.py also integrates mean and variance, checks controls and halves step size; plot.py draws densities from saved results. Figures are calculated; the cover is illustrative. This is AI-assisted educational analysis, not original research or an EL-AI experimental result.

T = 100.0
for gamma in (0, 1, 3):
    a0 = gamma + (1-gamma)/4
    aT = gamma/(1+T) + (1-gamma)/(4+T)
    static_variance = 1/a0
    transported_variance = (1/aT)*(1/(1+T))**gamma*(4/(4+T))**(1-gamma)
    print(gamma, static_variance, transported_variance)
# Exact variance of the specified Gaussian ODE, not an image benchmark.

Code, data, and instructions · JSON. Educational calculations executed with Python 3.14.0; figures with Matplotlib 3.11.2. AI-assisted analysis, without claiming peer review or human review. Original illustrative ImageGen cover: it does not document EL-AI people, premises, or installations. Sources accessed 29 September 2026.