Skip to content

Paired comparisons

Use common random numbers in your simulator: replicate index rep must identify the same random realization across configurations. Matching replicate labels alone does not establish this scientific assumption. Keep raw replicated grid results for these comparisons; aggregated means cannot recover paired draws.

paired_difference() estimates the mean difference A minus B and a bootstrap or paired-t interval. paired_rank() compares designs with a reference and applies Bonferroni adjustment by default. These comparisons address one observable; they do not replace multi-objective Pareto analysis.

trade_study.PairedDifference(design_a, design_b, observable, mean, lower, upper, n_pairs, confidence, method) dataclass

Mean per-replicate difference a - b with a confidence interval.

Attributes:

Name Type Description
design_a dict[str, Any]

Config of the first design.

design_b dict[str, Any]

Config of the second design.

observable str

Observable compared.

mean float

Mean of the per-replicate differences.

lower float

Lower confidence bound.

upper float

Upper confidence bound.

n_pairs int

Replicates with a finite value for both designs.

confidence float

Confidence level of the interval.

method str

"bootstrap" (percentile) or "t" (paired t).

trade_study.paired_difference(results, design_a, design_b, observable, *, method='bootstrap', confidence=0.95, n_boot=2000, seed=0)

Estimate the mean difference a - b over shared replicates.

Parameters:

Name Type Description Default
results ResultsTable

Per-replicate results; every row needs metadata["rep"].

required
design_a Design

First design, as a config subset that identifies one design point or as its metadata["design_point"] index.

required
design_b Design

Second design, specified the same way.

required
observable str

Name of the observable to compare.

required
method str

"bootstrap" for a percentile bootstrap over replicates, or "t" for a paired t interval (requires scipy).

'bootstrap'
confidence float

Two-sided confidence level.

0.95
n_boot int

Bootstrap resamples.

2000
seed int

Bootstrap seed.

0

Returns:

Type Description
PairedDifference

The paired difference and its interval. Replicates where either

PairedDifference

design's value is non-finite are dropped and not counted in

PairedDifference

n_pairs.

Raises:

Type Description
ValueError

If the designs' replicate sets differ, a design repeats a replicate, fewer than two finite pairs remain, or method is unknown.

Source code in src/trade_study/paired.py
def paired_difference(  # ruff: ignore[too-many-arguments]
    results: ResultsTable,
    design_a: Design,
    design_b: Design,
    observable: str,
    *,
    method: str = "bootstrap",
    confidence: float = 0.95,
    n_boot: int = 2000,
    seed: int = 0,
) -> PairedDifference:
    """Estimate the mean difference ``a - b`` over shared replicates.

    Args:
        results: Per-replicate results; every row needs ``metadata["rep"]``.
        design_a: First design, as a config subset that identifies one design
            point or as its ``metadata["design_point"]`` index.
        design_b: Second design, specified the same way.
        observable: Name of the observable to compare.
        method: ``"bootstrap"`` for a percentile bootstrap over replicates,
            or ``"t"`` for a paired t interval (requires scipy).
        confidence: Two-sided confidence level.
        n_boot: Bootstrap resamples.
        seed: Bootstrap seed.

    Returns:
        The paired difference and its interval. Replicates where either
        design's value is non-finite are dropped and not counted in
        ``n_pairs``.

    Raises:
        ValueError: If the designs' replicate sets differ, a design repeats a
            replicate, fewer than two finite pairs remain, or ``method`` is
            unknown.
    """
    if method not in METHODS:
        msg = f"method must be one of {METHODS}, got {method!r}"
        raise ValueError(msg)
    rows_a, config_a = _select(results, design_a)
    rows_b, config_b = _select(results, design_b)
    diff = _paired_differences(
        _column(results, observable),
        _by_rep(results, rows_a, config_a),
        _by_rep(results, rows_b, config_b),
    )
    lower, upper = (
        _bootstrap_interval(diff, confidence, n_boot, seed)
        if method == "bootstrap"
        else _t_interval(diff, confidence)
    )
    return PairedDifference(
        design_a=config_a,
        design_b=config_b,
        observable=observable,
        mean=float(diff.mean()),
        lower=lower,
        upper=upper,
        n_pairs=len(diff),
        confidence=confidence,
        method=method,
    )

trade_study.paired_rank(results, observable, reference, *, maximize=False, adjust='bonferroni', method='bootstrap', confidence=0.95, n_boot=2000, seed=0)

Compare every design with a reference design, best first.

Each entry is the paired difference design - reference. With adjust="bonferroni" each interval uses confidence 1 - (1 - confidence) / m for m comparisons, so all intervals hold jointly at confidence; with adjust=None they hold one at a time.

Parameters:

Name Type Description Default
results ResultsTable

Per-replicate results; every row needs metadata["rep"].

required
observable str

Name of the observable to compare.

required
reference Design

Reference design (config subset or design-point index).

required
maximize bool

Whether larger values are better (sets the ordering).

False
adjust str | None

"bonferroni" or None.

'bonferroni'
method str

Interval method, as in :func:paired_difference.

'bootstrap'
confidence float

Joint (adjusted) or per-comparison confidence level.

0.95
n_boot int

Bootstrap resamples.

2000
seed int

Bootstrap seed.

0

Returns:

Type Description
list[PairedDifference]

One paired difference per non-reference design, ordered from the

list[PairedDifference]

most improved on the reference to the least.

Raises:

Type Description
ValueError

If adjust is unknown.

Source code in src/trade_study/paired.py
def paired_rank(  # ruff: ignore[too-many-arguments]
    results: ResultsTable,
    observable: str,
    reference: Design,
    *,
    maximize: bool = False,
    adjust: str | None = "bonferroni",
    method: str = "bootstrap",
    confidence: float = 0.95,
    n_boot: int = 2000,
    seed: int = 0,
) -> list[PairedDifference]:
    """Compare every design with a reference design, best first.

    Each entry is the paired difference ``design - reference``. With
    ``adjust="bonferroni"`` each interval uses confidence
    ``1 - (1 - confidence) / m`` for ``m`` comparisons, so all intervals hold
    jointly at ``confidence``; with ``adjust=None`` they hold one at a time.

    Args:
        results: Per-replicate results; every row needs ``metadata["rep"]``.
        observable: Name of the observable to compare.
        reference: Reference design (config subset or design-point index).
        maximize: Whether larger values are better (sets the ordering).
        adjust: ``"bonferroni"`` or ``None``.
        method: Interval method, as in :func:`paired_difference`.
        confidence: Joint (adjusted) or per-comparison confidence level.
        n_boot: Bootstrap resamples.
        seed: Bootstrap seed.

    Returns:
        One paired difference per non-reference design, ordered from the
        most improved on the reference to the least.

    Raises:
        ValueError: If ``adjust`` is unknown.
    """
    if adjust not in {"bonferroni", None}:
        msg = f"adjust must be 'bonferroni' or None, got {adjust!r}"
        raise ValueError(msg)
    _rows, reference_config = _select(results, reference)
    others = [
        config
        for config in _design_configs(results)
        if _key(config) != _key(reference_config)
    ]
    level = confidence
    if adjust == "bonferroni" and others:
        level = 1 - (1 - confidence) / len(others)
    ranked = [
        paired_difference(
            results,
            config,
            reference_config,
            observable,
            method=method,
            confidence=level,
            n_boot=n_boot,
            seed=seed,
        )
        for config in others
    ]
    return sorted(ranked, key=lambda d: -d.mean if maximize else d.mean)