Skip to content

Introduce model_compare() for model comparison with support of new predictive measures - #380

Open
florence-bockting wants to merge 88 commits into
pred_measurefrom
integrate-loo_compare
Open

florence-bockting wants to merge 88 commits into
pred_measurefrom
integrate-loo_compare

Conversation

@florence-bockting

@florence-bockting florence-bockting commented Jul 7, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #220

Summary

This PR adds model_compare(). The function compares models on all predictive measures of the *_pred_measure() API (#363).

  • model_compare() replaces loo_compare().
  • loo_compare() is deprecated. It warns once per session. It still compares "loo", "waic", and "kfold" objects on ELPD.
  • loo_compare methods in other packages (e.g. loo_compare.brmsfit) still dispatch.
  • model_compare() accepts results from loo_pred_measure(), kfold_pred_measure(), test_pred_measure(), and
    insample_pred_measure().
  • For each measure that all models share, the function computes the paired difference and its standard error.
  • Each measure uses its own best model as the reference.
  • The function flips loss measures (e.g. mse) to the utility scale. A higher difference is then always better. print() marks each flipped measure.
  • The new custom_measure() defines a custom measure. It sets the name, the orientation, and the standard error of the difference.

The output for "loo", "waic", and "kfold" objects does not change.

The comment below lists all new features and functions.

The review decisions are in notes/design-discussions/model_compare.md.

Examples

@codecov-commenter

codecov-commenter commented Jul 7, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.40951% with 103 lines in your changes missing coverage. Please review.
✅ Project coverage is 91.68%. Comparing base (71b23c2) to head (0bd95d6).

Files with missing lines Patch % Lines
R/model_compare-print.R 87.36% 36 Missing ⚠️
R/model_compare-pred_measure.R 92.70% 28 Missing ⚠️
R/pred_measure-builtin.R 67.85% 27 Missing ⚠️
R/pred_measure-helpers.R 92.18% 5 Missing ⚠️
R/pred_measure-compute.R 95.89% 3 Missing ⚠️
R/loo_compare.R 75.00% 2 Missing ⚠️
R/model_compare.R 99.14% 2 Missing ⚠️
Additional details and impacted files
@@               Coverage Diff                @@
##           pred_measure     #380      +/-   ##
================================================
+ Coverage         91.29%   91.68%   +0.39%     
================================================
  Files                35       38       +3     
  Lines              4168     5049     +881     
================================================
+ Hits               3805     4629     +824     
- Misses              363      420      +57     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Jul 7, 2026 •

Copy link
Copy Markdown

This is how benchmark results would change (along with a 95% confidence interval in relative change) if 0bd95d6 is merged into pred_measure:

  • ✔️loo_function: 1.52s -> 1.51s [-1.3%, +0.58%]
  • ✔️loo_matrix: 1.38s -> 1.38s [-0.48%, +0.39%]
    Further explanation regarding interpretation and methodology can be found in the documentation.

florence-bockting and others added 22 commits July 8, 2026 10:49
…nals

Re-derives the file split on top of the current branch rather than merging
the earlier WIP, which git resolved into eight duplicate definitions.

model_compare.R (2082 lines) is split by concern into model_compare.R (the
generic, default method, ordering and diagnostics), model_compare-pred_measure.R
(the multi-measure path, rank resolution, standard errors) and
model_compare-print.R (print.compare.loo and its table helpers). Every moved
body is byte-identical to its previous version.

loo_compare.R keeps only the forwarding generic and its methods. The copies of
elpd_diffs, se_elpd_diff, find_model_names, middle_idx, order_stat_heuristic,
diag_elpd, diag_diff and print.compare.loo it also held were dead: alphabetical
collation meant model_compare.R won. diag_diff and print.compare.loo had
diverged, so the dead copies were also wrong. loo_compare_checks,
loo_compare_matrix, loo_compare_order and loo_order_stat_check are removed for
the same reason.

Warning helpers renamed to the package convention: .warn_insample_compare,
.warn_kfold_K_mismatch and .warn_omitted_compare_measures become
throw_*_warning(); .inform_compare_sign_conversion loses its dot prefix.

Adds loo_compare.psis_loo_ss_list so subsampling objects keep dispatching under
the old name, and fixes the undefined `loo3` in the loo_compare examples.

NAMESPACE and man/ still need regenerating with roxygen2 8.0.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The refactor copied an older print.compare.loo into model_compare.R, which
alphabetical collation made the live one, so the simplify argument added later
at a user's request stopped existing. Because print.compare.loo takes `...`,
`print(comp, simplify = FALSE)` was silently swallowed rather than erroring,
and the snapshot tests that would have caught it skip unless NOT_CRAN is set.

Merges the two: keeps the pred_measure dispatch, the compare_ref_model message
and .print_compare_diag_message() from the model_compare.R version, and restores
the flexible column selection and simplify argument from the loo_compare.R one.

test_compare.R with NOT_CRAN=true: 355 pass, 0 fail (was 352 pass, 3 fail on
8897e1f and on the split commit).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@florence-bockting

florence-bockting commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Overview of changes

Features

Comparison

  • Compares all pred_measure types. All models in one call must use the same evaluation source. Mixed sources cause an error.
  • Matches measures on bare names. rmse_loo and rmse_kfold both become rmse.
  • Uses the best model of each measure as the reference for that measure. The attribute compare_reference records each reference.
  • Rows sort by elpd when all models share it. Otherwise they sort by the first shared measure in alphabetical order.
  • Accepts named models: model_compare(A = m1, B = m2).
  • Still supports subsampled LOO (psis_loo_ss).

Standard error of differences

  • "sum" and "mean": from paired pointwise differences.
  • "measure_specific": special formulas for rmse, r2, and bacc.
  • "custom": set with attr(fun, "measure_se_diff") or custom_measure(). If a custom measure sets nothing, the SE is NA. A message then tells the user.

Checks

Condition Output
Models use different measure sets warning (lists omitted measures)
Models disagree on loss or se_diff_fun error
Models use different y warning
k-fold models use different K or folds warning
In-sample comparison warning (optimistic bias)
More than 11 models warning (order-statistic check)
loo_moment_match() or reloo() with a measure other than elpd, mlpd, ic warning
print.compare.loo() when measures is given for loo objects warning

Print method (print.compare.loo())

  • measures: NULL (first measure), "all", or a character vector.
  • simplify = FALSE: also shows the per-model estimate and SE.
  • digits = NULL: sets the digits for each measure.
  • Prints the reference model for each measure table. With more than four measures, the print names only the first reference, e.g. (elpd: m2, ...)
  • Prints the PSIS Pareto k diagnostics once per model, above the tables.
  • Marks each loss measure with sign flipped in the table header. A note below the tables names these measures.
  • Shows the diagnostic glossary only when the output includes an ELPD measure.
  • Keeps each output line at 80 characters or fewer.

User-facing functions

Function Status
model_compare() (methods default, psis_loo_ss_list) new
custom_measure() new
print.compare.loo() new arguments measures, simplify
loo_compare() deprecated
print.compare.loo_ss() removed

Internal functions

R/model_compare.R

  • New: .model_compare_inputs(), .measure_ref_model(), .model_compare_estimates_table(), throw_kfold_K_mismatch_warning(), throw_kfold_folds_mismatch_warning()
  • Renamed from loo_compare_*: model_compare_checks(), model_compare_matrix(), model_compare_order(), model_order_stat_check()
  • Moved from R/loo_compare.R: elpd_diffs(), se_elpd_diff(), find_model_names(), middle_idx(), order_stat_heuristic(), diag_elpd(), diag_diff()

R/model_compare-pred_measure.R (new file)

  • Dispatch and checks: is.pred_measure(), is.loo_pred_measure(), compare_pred_measure(), .compare_source(), .compare_metadata_check()
  • Measure sets: .compare_measures(), .compare_pointwise_cols(), .resolve_rank_measure(), .is_elpd_measure(), .get_measure_info()
  • Orientation: .builtin_loss_measures(), .measure_is_loss(), .compare_sign_converted_measures()
  • SE of differences: .measure_pointwise_diff_method(), .check_declared_aggregation(), .measure_se_diff_fun(), .check_se_diff_value(), .resolve_custom_se_diffs(), .se_diff_input(), .validate_se_diff(), .pair_measure_stats()
  • Conditions: throw_insample_compare_warning(), throw_omitted_compare_measures_warning(), inform_missing_custom_se_diff()

R/model_compare-print.R (new file)

  • .print_compare_pred_measure(), .print_compare_measure_table(), .print_psis_diag_block(), .print_compare_diag_message(), .compare_reference_line(), .compare_source_phrase(), .parse_diag_psis(),
    .compare_suffix(), .display_name(), .n_of_models(), .cat_wrapped(), .print_sign_flip_note()

R/model_compare.psis_loo_ss_list.R (renamed from loo_compare.psis_loo_ss_list.R)

  • model_compare_ss(), model_compare_ss_naive(), model_compare_ss_diff()

R/pred_measure-builtin.R

  • .se_diff_rmse(), .se_diff_r2(), .se_diff_bacc(), .se_r2_delta(), .measure_info()

R/pred_measure-compute.R

  • .detect_posthoc(), .warn_posthoc()

R/print.R

  • .measure_digits(), .se_digits(), .resolve_digits(), .format_estimates()

Florence Bockting added 2 commits September 28, 2026 11:24
…_compare

# Conflicts:
#	R/pred_measure-compute.R
#	R/pred_measure-helpers.R
#	R/pred_measure.R
#	man/insample_pred_measure.Rd
#	man/pred_measure_params.Rd
@florence-bockting
florence-bockting marked this pull request as ready for review September 28, 2026 09:00

@jgabry jgabry left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @florence-bockting, this is great! Here's a first round of review comments, but I probably missed some things. I should probably do another round, but we can start with these first and then I'll review again

Comment thread tests/testthat/data-for-tests/test_data_roaches_compare.Rds
Comment thread R/model_compare-pred_measure.R
Comment thread R/model_compare-pred_measure.R Outdated
Comment thread R/model_compare.R
Comment thread R/model_compare.R Outdated
Comment thread vignettes/articles-online-only/model-comparison.Rmd Outdated
Comment thread vignettes/articles-online-only/model-comparison.Rmd Outdated
Comment thread R/pred_measure.R
Comment thread R/pred_measure.R Outdated
Comment thread R/pred_measure.R Outdated
@florence-bockting

Copy link
Copy Markdown
Contributor Author

Thank you @jgabry for the great review. I went now through all you comments and updated the code base correspondingly.

@jgabry jgabry left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes look good. I made a bunch more comments but they're mostly more cases where attributes are referred to that we should rewrite. The only other thing was about one test.

After fixing those I think we can merge this into the other branch. Maybe I'll notice more things when I review the other branch when it's ready.

"Not all kfold objects have the same K value"
)

test_that("print warns that `measures` is ignored for 'loo' comparisons", {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This test_that() ended up inside the "model_compare throws appropriate warnings" one (the one starting at line 1331). It still runs, but it should probably be moved out from inside the other test

Comment thread R/loo-glossary.R
#' `r2` it is the trivariate analogue, which additionally propagates the
#' uncertainty in the baseline `MSE(y)` shared by both models.
#' * For custom measures it comes from the measure's own
#' `attr(my_fun, "measure_se_diff")` declaration, set with

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's avoid mentioning the attribute here and just say it comes from what the user passes to custom_measure()

Comment thread R/loo-glossary.R
#' * `se_diff_fun`: for built-in measures with
#' `diff_method = "measure_specific"`, the name of the built-in implementation
#' used. For custom measures, whatever the measure declared in
#' `attr(my_fun, "measure_se_diff")`; absent when it declared nothing.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another attributes mention that we can rephrase to avoid mentioning the attribute itself

Comment thread R/loo-glossary.R
#' [custom_measure()]. With `loss = TRUE` lower values are better; without it
#' they are treated as utilities (see [insample_pred_measure()]).
#' [model_compare()] requires all models to provide matching `measure_info` for
#' each shared measure; a mismatched `measure_loss` or `measure_se_diff`

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These are the attributes right? I think we could rephrase this too

Comment thread R/model_compare.R
#' pointwise contributions (`r2`, `rmse`, `bacc`), so the measure supplies
#' its own standard error of the difference.
#' * `"custom"`: the standard error comes from the measure's own
#' `attr(my_fun, "measure_se_diff")` declaration, set with

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another attributes mention that we can rephrase to avoid mentioning the attribute itself

"Models disagree on `measure_info` for measure '",
bare,
"'. For a custom measure, ensure all models use the same ",
"`measure_loss` and `measure_se_diff` declarations.",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another attributes mention that we can rephrase to avoid mentioning the attributes themselves

}
if (!ok) {
warning(
"`measure_se_diff = \"", method, "\"` was declared for measure '",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another attributes mention that we can rephrase to avoid mentioning the attribute itself

Comment thread R/pred_measure-helpers.R
!identical(attr_name, name)) {
cli::cli_warn(c(
"Custom measure named {.val {name}} in {.arg measure} also has",
"{.code attr(fun, \"measure_name\") = {.val {attr_name}}}.",

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Another attributes mention that we can rephrase to avoid mentioning the attribute itself

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants