Package {seqbench}


Title: Anytime-Valid Sequential Benchmarking of Algorithms
Version: 0.1.0
Description: Sequential, anytime-valid comparison of algorithms on paired losses with a practical-equivalence margin. Confidence sequences for the mean paired difference, superiority/equivalence/continue decisions, cost accounting and auditable trajectories.
License: GPL (≥ 3)
URL: https://github.com/castlaboratory/seqbench
BugReports: https://github.com/castlaboratory/seqbench/issues
Encoding: UTF-8
RoxygenNote: 8.0.0
Depends: R (≥ 4.1)
Imports: cli (≥ 3.6.0), generics, ggplot2 (≥ 3.5.0), rlang (≥ 1.1.0), stats, tibble (≥ 3.2.0), utils
Suggests: knitr, rmarkdown, testthat (≥ 3.0.0)
Config/testthat/edition: 3
VignetteBuilder: knitr
NeedsCompilation: no
Packaged: 2026-09-28 13:26:41 UTC; leite
Author: André Leite [aut, cre], Raydonal Ospina [aut], Cristiano Ferraz [aut]
Maintainer: André Leite <leite@castlab.org>
Repository: CRAN
Date/Publication: 2026-10-08 11:20:02 UTC

seqbench: Anytime-Valid Sequential Benchmarking of Algorithms

Description

Sequential, anytime-valid comparison of algorithms on paired losses with a practical-equivalence margin. Confidence sequences for the mean paired difference, superiority/equivalence/continue decisions, cost accounting and auditable trajectories.

Author(s)

Maintainer: André Leite leite@castlab.org

Authors:

See Also

Useful links:


Plot the inferential trajectory of a comparison

Description

Running-intersection confidence sequence for the mean paired difference against the number of instances, with the equivalence band ⁠[-margin, margin]⁠ and the zero line.

Usage

autoplot.seqbench_comparison(object, ...)

Arguments

object

A seqbench_comparison with at least one instance.

...

Unused.

Value

A ggplot object.

Examples

design <- comparison_design(margin = 0.05, bounds = c(0, 1), boundary = "hoeffding")
set.seed(2)
losses <- data.frame(instance = 1:40, loss_a = runif(40, 0, .5), loss_b = runif(40, .2, .7))
state <- update_comparison(initialize_comparison(design), losses)
ggplot2::autoplot(state)

Boundary kernels (internal API)

Description

Low-level state machines behind the confidence sequences used by update_comparison(). Observations must already be rescaled to ⁠[0, 1]⁠. These functions are exported for reuse by sister packages; ordinary users should call comparison_design() and update_comparison() instead.

Usage

boundary_init(
  name = seqbench_boundaries(),
  alpha = 0.05,
  c = 0.5,
  theta = 0.5,
  grid = 1001L,
  refine = TRUE,
  sd_max01 = NULL
)

boundary_update(state, x)

boundary_interval(state, thresholds = NULL)

Arguments

name

One of seqbench_boundaries().

alpha

Miscoverage level in (0, 1).

c

Truncation constant for the empirical-Bernstein and betting bets, in (0, 1). Waudby-Smith & Ramdas recommend 1/2 or 3/4.

theta

Hedging weight on the "long" capital in the betting boundary, in (0, 1). Default 1/2.

grid

Number of grid points on ⁠[0, 1]⁠ for the betting boundary.

refine

Logical; refine the betting interval endpoints by bisection on the exact capital (default TRUE). Without refinement the endpoints are the outer boundaries of the grid cells that contain the true endpoints, which is conservative and keeps the coverage guarantee.

sd_max01

For "bernstein_declared": declared upper bound on the conditional standard deviation of the observations, on the ⁠[0, 1]⁠ scale (at most 1/2). The boundary is a predictable-plug-in Bennett confidence sequence: for ⁠X in [0, 1]⁠ with conditional mean mu and conditional variance at most sd_max01^2, exp(lambda (X - mu) - sd_max01^2 (e^lambda - 1 - lambda)) is a supermartingale for every predictable lambda >= 0 (Bennett's inequality), and the same holds for -(X - mu). Coverage is guaranteed only if the declared bound is true.

state

A boundary state returned by boundary_init() or boundary_update().

x

A single number in ⁠[0, 1]⁠.

thresholds

Optional numeric vector of decision thresholds on the ⁠[0, 1]⁠ scale. When given to boundary_interval() for the betting boundary, an endpoint is refined only if its grid cell contains one of them, which is the only case in which refinement can change a decision; this keeps the per-step cost linear in the grid size instead of linear in the number of observations. Decisions and stopping times are identical to full refinement; the reported running intersection can be up to one grid cell wider per side. NULL (default) refines both endpoints.

Value

boundary_init() and boundary_update() return a state object of class seqbench_boundary. boundary_interval() returns a named numeric vector c(estimate, lower, upper) on the ⁠[0, 1]⁠ scale. For the betting boundary, NA endpoints mean that the exact confidence set is empty (a possible event, of probability at most alpha under the assumptions); a nonempty set narrower than one grid cell is located by searching the exact capital, never reported as empty.

Examples

st <- boundary_init("hoeffding", alpha = 0.05)
for (x in c(0.2, 0.3, 0.25)) st <- boundary_update(st, x)
boundary_interval(st)

Design a sequential comparison of two algorithms

Description

Fixes, before any data is seen, everything the procedure needs: the error level, the margin of practical equivalence, the loss bounds, the boundary and the budget. The design is immutable; initialize_comparison() turns it into a state that update_comparison() advances.

Usage

comparison_design(
  alpha = 0.05,
  margin,
  bounds,
  boundary = "betting",
  paired = TRUE,
  cost_per_round = 1,
  budget = Inf,
  n_max = Inf,
  betting_c = 0.5,
  betting_theta = 0.5,
  betting_grid = 1001L,
  sd_max = NULL
)

Arguments

alpha

Miscoverage level of the confidence sequence, in (0, 1). The probability that the procedure ever issues a false declaration (of any of the three kinds) is at most alpha.

margin

Margin of practical equivalence delta > 0, in loss units. Algorithm A is declared relevantly better when the whole confidence sequence for mean(loss_a - loss_b) lies below -margin; B when it lies above margin; the two are equivalent when it lies inside ⁠[-margin, margin]⁠.

bounds

Known bounds c(lower, upper) of the losses of both algorithms. Paired differences then lie in ⁠[lower - upper, upper - lower]⁠. A loss outside the bounds is an error, never clipped.

boundary

One of seqbench_boundaries(). "betting" (default) is the hedged capital confidence sequence of Waudby-Smith & Ramdas (2024); "empirical_bernstein" and "hoeffding" are their conservative predictable-plug-in references; "bernstein_declared" is a predictable-plug-in Bennett confidence sequence that uses a declared upper bound sd_max on the standard deviation of the paired difference and is valid only if that bound holds. Its gain over the betting boundary is modest (a logarithmic factor in the declared variance): for bounded observations the Bennett bound keeps a range term of order log(2/alpha) / t whatever the variance, so the boundary is provided for completeness and for the package's cost study rather than as a shortcut; "naive_fixed" is a fixed-sample t interval recomputed at every step, not valid under optional stopping, provided only as a negative control for experiments.

paired

Must be TRUE. Unpaired designs are not implemented; the argument exists so that the limitation is explicit.

cost_per_round

Default cost of one ⁠(instance, seed)⁠ evaluation of both algorithms together, used when the data carry no cost column.

budget

Total cost after which the procedure stops with an inconclusive outcome. Inf for no cost limit.

n_max

Maximum number of instances. Inf for no limit.

betting_c, betting_theta, betting_grid

Tuning of the betting boundary: truncation constant of the bets (default 1/2, the value recommended by Waudby-Smith & Ramdas, who also suggest 3/4), hedging weight (default 1/2) and grid size on ⁠[0, 1]⁠ (default 1001). betting_c is also the truncation constant of the empirical-Bernstein boundary. Larger values bet more aggressively and stop earlier: in the package's simulation study (2,000 replications per cell) betting_c = 0.9 needed about 40% fewer instances than 1/2 at low noise and 20% fewer at high noise, with the largest observed coverage failure rising from 0.6% to 1.4% at alpha = 0.05; 0.99 gains little more. Coverage is guaranteed for any value in (0, 1).

sd_max

Required for boundary = "bernstein_declared": declared upper bound on the standard deviation of the per-instance (seed-averaged) paired difference, in loss units. This is an additional assumption (A5) recorded in the result contract.

Value

An object of class seqbench_design (a list).

References

Waudby-Smith, I. and Ramdas, A. (2024). Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B, 86(1), 1-27. doi:10.1093/jrsssb/qkad009

See Also

initialize_comparison(), planning_horizon()

Examples

design <- comparison_design(alpha = 0.05, margin = 0.02, bounds = c(0, 1))
design

Full result contract of a comparison

Description

Returns every element the protocol promises: estimand, estimate, current confidence sequence, required assumptions, diagnostics, decision and stopping reason, resources consumed, seeds and versions. summary() is an alias.

Usage

comparison_report(state)

## S3 method for class 'seqbench_comparison'
summary(object, ...)

Arguments

state

A seqbench_comparison.

object

A seqbench_comparison.

...

Unused.

Value

A list of class seqbench_report.

Examples

state <- initialize_comparison(comparison_design(margin = 0.02, bounds = c(0, 1)))
comparison_report(state)

Glance at a comparison

Description

Glance at a comparison

Usage

## S3 method for class 'seqbench_comparison'
glance(x, ...)

Arguments

x

A seqbench_comparison.

...

Unused.

Value

A one-row tibble: boundary, valid, alpha, margin, n_instances, n_evaluations, cost, estimate, lower, upper, decision, stopping_reason.

Examples

state <- initialize_comparison(comparison_design(margin = 0.02, bounds = c(0, 1)))
glance(state)

Start a sequential comparison

Description

Creates the state object that update_comparison() advances one instance at a time. Nothing is observed yet; the confidence sequence is ⁠[a, b]⁠, the full range of the paired difference.

Usage

initialize_comparison(design)

Arguments

design

A comparison_design().

Value

An object of class seqbench_comparison. Its main components are design, data (all rows seen so far), trajectory (one row per instance: estimate, confidence sequence, running intersection, cost and decision at that time), decision, stopping_reason, diagnostics and meta (versions, timestamps). Use tidy() for the trajectory, glance() for a one-row summary and comparison_report() for the full contract.

Examples

design <- comparison_design(margin = 0.02, bounds = c(0, 1))
state <- initialize_comparison(design)
state

Planning horizon for a distance to the decision threshold

Description

Number of instances after which the confidence sequence is expected to have declared the true hypothesis, when the true mean paired difference is at distance distance from the nearest threshold -margin or margin. Two modes:

Usage

planning_horizon(design, distance, sd = NULL, max_t = 1e+06)

Arguments

design

A comparison_design().

distance

Positive distance ⁠|mu| - margin⁠ (superiority) or ⁠margin - |mu|⁠ (equivalence), in loss units.

sd

Optional standard deviation of the paired difference, loss units.

max_t

Search limit.

Details

Both charge the full half-width twice (worst-case position of the centre); typical stopping times are about half of the returned value, as measured in the package's pilot study. Both are computable before any data are collected and serve to size budget and n_max.

Value

An integer number of instances, or Inf if not reached by max_t.

Examples

design <- comparison_design(margin = 0.02, bounds = c(0, 1))
planning_horizon(design, distance = 0.05)             # guaranteed, Hoeffding
planning_horizon(design, distance = 0.05, sd = 0.1)   # variance-based approximation

Objects exported from other packages

Description

These objects are imported from other packages. Follow the links below to see their documentation.

generics

glance(), tidy()


Names of the available boundaries

Description

Names of the available boundaries

Usage

seqbench_boundaries()

Value

A character vector. "betting" is the default in comparison_design(); "hoeffding" and "empirical_bernstein" are conservative references; "bernstein_declared" requires a declared upper bound on the standard deviation of the paired difference and is valid only if that bound holds; "naive_fixed" is an invalid negative control for experiments.

Examples

seqbench_boundaries()

Current decision of a comparison

Description

Current decision of a comparison

Usage

stopping_decision(state)

Arguments

state

A seqbench_comparison.

Value

A character scalar: "A" (A relevantly better), "B", "equivalent", "continue" or "inconclusive", with attribute reason ("declaration", "budget", "n_max", "cs_empty" or NA).

Examples

state <- initialize_comparison(comparison_design(margin = 0.02, bounds = c(0, 1)))
stopping_decision(state)

Tidy the trajectory of a comparison

Description

One row per instance with the estimate, the confidence sequence at that time, its running intersection, cumulative cost and the decision the procedure would have taken there.

Usage

## S3 method for class 'seqbench_comparison'
tidy(x, ...)

Arguments

x

A seqbench_comparison.

...

Unused.

Value

A tibble with columns t, instance, n_seeds, x, estimate, lower, upper, lower_running, upper_running, cost, cost_cum, decision.

Examples

design <- comparison_design(margin = 0.05, bounds = c(0, 1), boundary = "hoeffding")
state <- update_comparison(initialize_comparison(design),
  data.frame(instance = 1:5, loss_a = c(.1, .2, .1, .3, .2), loss_b = c(.5, .6, .4, .7, .5)))
tidy(state)

Update a comparison with new paired losses

Description

Absorbs the losses of one or more new instances, in the order given, and advances the confidence sequence and the decision after each instance. The sequential unit is the instance: all rows of an instance (its seeds) are averaged into one observation. Updating stops at the first instance that triggers a terminal decision or exhausts the budget; rows after it are not consumed and a warning says so.

Usage

update_comparison(state, losses)

Arguments

state

A seqbench_comparison from initialize_comparison().

losses

A data frame with columns instance, loss_a, loss_b, and optionally seed (default 1) and cost (default design$cost_per_round per row). Losses must lie within the design bounds; instances must not have been seen before.

Value

The updated seqbench_comparison.

Examples

design <- comparison_design(margin = 0.05, bounds = c(0, 1), boundary = "hoeffding")
state <- initialize_comparison(design)
set.seed(1)
losses <- data.frame(instance = 1:50, loss_a = runif(50, 0, 0.5),
                     loss_b = runif(50, 0.3, 0.8))
state <- update_comparison(state, losses)
stopping_decision(state)