Paired comparisons¶
Use common random numbers in your simulator: replicate index rep must identify
the same random realization across configurations. Matching replicate labels
alone does not establish this scientific assumption. Keep raw replicated grid
results for these comparisons; aggregated means cannot recover paired draws.
paired_difference() estimates the mean difference A minus B and a bootstrap or
paired-t interval. paired_rank() compares designs with a reference and applies
Bonferroni adjustment by default. These comparisons address one observable;
they do not replace multi-objective Pareto analysis.
trade_study.PairedDifference(design_a, design_b, observable, mean, lower, upper, n_pairs, confidence, method)
dataclass
¶
Mean per-replicate difference a - b with a confidence interval.
Attributes:
| Name | Type | Description |
|---|---|---|
design_a |
dict[str, Any]
|
Config of the first design. |
design_b |
dict[str, Any]
|
Config of the second design. |
observable |
str
|
Observable compared. |
mean |
float
|
Mean of the per-replicate differences. |
lower |
float
|
Lower confidence bound. |
upper |
float
|
Upper confidence bound. |
n_pairs |
int
|
Replicates with a finite value for both designs. |
confidence |
float
|
Confidence level of the interval. |
method |
str
|
|
trade_study.paired_difference(results, design_a, design_b, observable, *, method='bootstrap', confidence=0.95, n_boot=2000, seed=0)
¶
Estimate the mean difference a - b over shared replicates.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results
|
ResultsTable
|
Per-replicate results; every row needs |
required |
design_a
|
Design
|
First design, as a config subset that identifies one design
point or as its |
required |
design_b
|
Design
|
Second design, specified the same way. |
required |
observable
|
str
|
Name of the observable to compare. |
required |
method
|
str
|
|
'bootstrap'
|
confidence
|
float
|
Two-sided confidence level. |
0.95
|
n_boot
|
int
|
Bootstrap resamples. |
2000
|
seed
|
int
|
Bootstrap seed. |
0
|
Returns:
| Type | Description |
|---|---|
PairedDifference
|
The paired difference and its interval. Replicates where either |
PairedDifference
|
design's value is non-finite are dropped and not counted in |
PairedDifference
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If the designs' replicate sets differ, a design repeats a
replicate, fewer than two finite pairs remain, or |
Source code in src/trade_study/paired.py
trade_study.paired_rank(results, observable, reference, *, maximize=False, adjust='bonferroni', method='bootstrap', confidence=0.95, n_boot=2000, seed=0)
¶
Compare every design with a reference design, best first.
Each entry is the paired difference design - reference. With
adjust="bonferroni" each interval uses confidence
1 - (1 - confidence) / m for m comparisons, so all intervals hold
jointly at confidence; with adjust=None they hold one at a time.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results
|
ResultsTable
|
Per-replicate results; every row needs |
required |
observable
|
str
|
Name of the observable to compare. |
required |
reference
|
Design
|
Reference design (config subset or design-point index). |
required |
maximize
|
bool
|
Whether larger values are better (sets the ordering). |
False
|
adjust
|
str | None
|
|
'bonferroni'
|
method
|
str
|
Interval method, as in :func: |
'bootstrap'
|
confidence
|
float
|
Joint (adjusted) or per-comparison confidence level. |
0.95
|
n_boot
|
int
|
Bootstrap resamples. |
2000
|
seed
|
int
|
Bootstrap seed. |
0
|
Returns:
| Type | Description |
|---|---|
list[PairedDifference]
|
One paired difference per non-reference design, ordered from the |
list[PairedDifference]
|
most improved on the reference to the least. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |