| Title: | Anytime-Valid Sequential Benchmarking of Algorithms |
| Version: | 0.1.0 |
| Description: | Sequential, anytime-valid comparison of algorithms on paired losses with a practical-equivalence margin. Confidence sequences for the mean paired difference, superiority/equivalence/continue decisions, cost accounting and auditable trajectories. |
| License: | GPL (≥ 3) |
| URL: | https://github.com/castlaboratory/seqbench |
| BugReports: | https://github.com/castlaboratory/seqbench/issues |
| Encoding: | UTF-8 |
| RoxygenNote: | 8.0.0 |
| Depends: | R (≥ 4.1) |
| Imports: | cli (≥ 3.6.0), generics, ggplot2 (≥ 3.5.0), rlang (≥ 1.1.0), stats, tibble (≥ 3.2.0), utils |
| Suggests: | knitr, rmarkdown, testthat (≥ 3.0.0) |
| Config/testthat/edition: | 3 |
| VignetteBuilder: | knitr |
| NeedsCompilation: | no |
| Packaged: | 2026-09-28 13:26:41 UTC; leite |
| Author: | André Leite [aut, cre], Raydonal Ospina [aut], Cristiano Ferraz [aut] |
| Maintainer: | André Leite <leite@castlab.org> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-08 11:20:02 UTC |
seqbench: Anytime-Valid Sequential Benchmarking of Algorithms
Description
Sequential, anytime-valid comparison of algorithms on paired losses with a practical-equivalence margin. Confidence sequences for the mean paired difference, superiority/equivalence/continue decisions, cost accounting and auditable trajectories.
Author(s)
Maintainer: André Leite leite@castlab.org
Authors:
André Leite leite@castlab.org
Raydonal Ospina raydonal@castlab.org
Cristiano Ferraz cferraz@castlab.org
See Also
Useful links:
Report bugs at https://github.com/castlaboratory/seqbench/issues
Plot the inferential trajectory of a comparison
Description
Running-intersection confidence sequence for the mean paired difference
against the number of instances, with the equivalence band
[-margin, margin] and the zero line.
Usage
autoplot.seqbench_comparison(object, ...)
Arguments
object |
A |
... |
Unused. |
Value
A ggplot object.
Examples
design <- comparison_design(margin = 0.05, bounds = c(0, 1), boundary = "hoeffding")
set.seed(2)
losses <- data.frame(instance = 1:40, loss_a = runif(40, 0, .5), loss_b = runif(40, .2, .7))
state <- update_comparison(initialize_comparison(design), losses)
ggplot2::autoplot(state)
Boundary kernels (internal API)
Description
Low-level state machines behind the confidence sequences used by
update_comparison(). Observations must already be rescaled to [0, 1].
These functions are exported for reuse by sister packages; ordinary users
should call comparison_design() and update_comparison() instead.
Usage
boundary_init(
name = seqbench_boundaries(),
alpha = 0.05,
c = 0.5,
theta = 0.5,
grid = 1001L,
refine = TRUE,
sd_max01 = NULL
)
boundary_update(state, x)
boundary_interval(state, thresholds = NULL)
Arguments
name |
One of |
alpha |
Miscoverage level in (0, 1). |
c |
Truncation constant for the empirical-Bernstein and betting bets, in (0, 1). Waudby-Smith & Ramdas recommend 1/2 or 3/4. |
theta |
Hedging weight on the "long" capital in the betting boundary, in (0, 1). Default 1/2. |
grid |
Number of grid points on |
refine |
Logical; refine the betting interval endpoints by bisection on
the exact capital (default |
sd_max01 |
For |
state |
A boundary state returned by |
x |
A single number in |
thresholds |
Optional numeric vector of decision thresholds on the
|
Value
boundary_init() and boundary_update() return a state object of
class seqbench_boundary. boundary_interval() returns a named numeric
vector c(estimate, lower, upper) on the [0, 1] scale. For the betting
boundary, NA endpoints mean that the exact confidence set is empty (a
possible event, of probability at most alpha under the assumptions); a
nonempty set narrower than one grid cell is located by searching the exact
capital, never reported as empty.
Examples
st <- boundary_init("hoeffding", alpha = 0.05)
for (x in c(0.2, 0.3, 0.25)) st <- boundary_update(st, x)
boundary_interval(st)
Design a sequential comparison of two algorithms
Description
Fixes, before any data is seen, everything the procedure needs: the error
level, the margin of practical equivalence, the loss bounds, the boundary and
the budget. The design is immutable; initialize_comparison() turns it into
a state that update_comparison() advances.
Usage
comparison_design(
alpha = 0.05,
margin,
bounds,
boundary = "betting",
paired = TRUE,
cost_per_round = 1,
budget = Inf,
n_max = Inf,
betting_c = 0.5,
betting_theta = 0.5,
betting_grid = 1001L,
sd_max = NULL
)
Arguments
alpha |
Miscoverage level of the confidence sequence, in (0, 1). The
probability that the procedure ever issues a false declaration (of any of
the three kinds) is at most |
margin |
Margin of practical equivalence |
bounds |
Known bounds |
boundary |
One of |
paired |
Must be |
cost_per_round |
Default cost of one |
budget |
Total cost after which the procedure stops with an
inconclusive outcome. |
n_max |
Maximum number of instances. |
betting_c, betting_theta, betting_grid |
Tuning of the betting boundary:
truncation constant of the bets (default 1/2, the value recommended by
Waudby-Smith & Ramdas, who also suggest 3/4), hedging weight (default
1/2) and grid size on |
sd_max |
Required for |
Value
An object of class seqbench_design (a list).
References
Waudby-Smith, I. and Ramdas, A. (2024). Estimating means of bounded random variables by betting. Journal of the Royal Statistical Society Series B, 86(1), 1-27. doi:10.1093/jrsssb/qkad009
See Also
initialize_comparison(), planning_horizon()
Examples
design <- comparison_design(alpha = 0.05, margin = 0.02, bounds = c(0, 1))
design
Full result contract of a comparison
Description
Returns every element the protocol promises: estimand, estimate, current
confidence sequence, required assumptions, diagnostics, decision and
stopping reason, resources consumed, seeds and versions. summary() is an
alias.
Usage
comparison_report(state)
## S3 method for class 'seqbench_comparison'
summary(object, ...)
Arguments
state |
A |
object |
A |
... |
Unused. |
Value
A list of class seqbench_report.
Examples
state <- initialize_comparison(comparison_design(margin = 0.02, bounds = c(0, 1)))
comparison_report(state)
Glance at a comparison
Description
Glance at a comparison
Usage
## S3 method for class 'seqbench_comparison'
glance(x, ...)
Arguments
x |
A |
... |
Unused. |
Value
A one-row tibble: boundary, valid, alpha, margin,
n_instances, n_evaluations, cost, estimate, lower, upper,
decision, stopping_reason.
Examples
state <- initialize_comparison(comparison_design(margin = 0.02, bounds = c(0, 1)))
glance(state)
Start a sequential comparison
Description
Creates the state object that update_comparison() advances one instance
at a time. Nothing is observed yet; the confidence sequence is [a, b], the
full range of the paired difference.
Usage
initialize_comparison(design)
Arguments
design |
Value
An object of class seqbench_comparison. Its main components are
design, data (all rows seen so far), trajectory (one row per
instance: estimate, confidence sequence, running intersection, cost and
decision at that time), decision, stopping_reason, diagnostics and
meta (versions, timestamps). Use tidy() for the trajectory, glance()
for a one-row summary and comparison_report() for the full contract.
Examples
design <- comparison_design(margin = 0.02, bounds = c(0, 1))
state <- initialize_comparison(design)
state
Planning horizon for a distance to the decision threshold
Description
Number of instances after which the confidence sequence is expected to have
declared the true hypothesis, when the true mean paired difference is at
distance distance from the nearest threshold -margin or margin. Two
modes:
Usage
planning_horizon(design, distance, sd = NULL, max_t = 1e+06)
Arguments
design |
|
distance |
Positive distance |
sd |
Optional standard deviation of the paired difference, loss units. |
max_t |
Search limit. |
Details
-
sd = NULL(default): the guaranteed horizont*(Delta) = min{t : 2 w_t < |Delta|}of the predictable-plug-in Hoeffding boundary, whose half-widthw_tis deterministic. On the coverage event (probability at least1 - alpha) the Hoeffding comparison has stopped with the correct declaration by then (theory note, Prop. 2). It is distribution-free and, for small margins, very conservative: it can exceed what the default betting boundary needs by two or three orders of magnitude. -
sdgiven (standard deviation of the paired difference, loss units): an approximation that plugssdinto the empirical-Bernstein half-width in place of the running variance estimate. It is not a bound; it is the order of magnitude a variance-adaptive boundary needs, and the betting boundary is usually faster still. Use a pilot or a conservative guess forsd.
Both charge the full half-width twice (worst-case position of the centre);
typical stopping times are about half of the returned value, as measured in
the package's pilot study. Both are computable before any data are collected
and serve to size budget and n_max.
Value
An integer number of instances, or Inf if not reached by max_t.
Examples
design <- comparison_design(margin = 0.02, bounds = c(0, 1))
planning_horizon(design, distance = 0.05) # guaranteed, Hoeffding
planning_horizon(design, distance = 0.05, sd = 0.1) # variance-based approximation
Objects exported from other packages
Description
These objects are imported from other packages. Follow the links below to see their documentation.
Names of the available boundaries
Description
Names of the available boundaries
Usage
seqbench_boundaries()
Value
A character vector. "betting" is the default in
comparison_design(); "hoeffding" and "empirical_bernstein" are
conservative references; "bernstein_declared" requires a declared upper
bound on the standard deviation of the paired difference and is valid only
if that bound holds; "naive_fixed" is an invalid negative control for
experiments.
Examples
seqbench_boundaries()
Current decision of a comparison
Description
Current decision of a comparison
Usage
stopping_decision(state)
Arguments
state |
A |
Value
A character scalar: "A" (A relevantly better), "B",
"equivalent", "continue" or "inconclusive", with attribute
reason ("declaration", "budget", "n_max", "cs_empty" or NA).
Examples
state <- initialize_comparison(comparison_design(margin = 0.02, bounds = c(0, 1)))
stopping_decision(state)
Tidy the trajectory of a comparison
Description
One row per instance with the estimate, the confidence sequence at that time, its running intersection, cumulative cost and the decision the procedure would have taken there.
Usage
## S3 method for class 'seqbench_comparison'
tidy(x, ...)
Arguments
x |
A |
... |
Unused. |
Value
A tibble with columns t, instance, n_seeds, x,
estimate, lower, upper, lower_running, upper_running, cost,
cost_cum, decision.
Examples
design <- comparison_design(margin = 0.05, bounds = c(0, 1), boundary = "hoeffding")
state <- update_comparison(initialize_comparison(design),
data.frame(instance = 1:5, loss_a = c(.1, .2, .1, .3, .2), loss_b = c(.5, .6, .4, .7, .5)))
tidy(state)
Update a comparison with new paired losses
Description
Absorbs the losses of one or more new instances, in the order given, and advances the confidence sequence and the decision after each instance. The sequential unit is the instance: all rows of an instance (its seeds) are averaged into one observation. Updating stops at the first instance that triggers a terminal decision or exhausts the budget; rows after it are not consumed and a warning says so.
Usage
update_comparison(state, losses)
Arguments
state |
A |
losses |
A data frame with columns |
Value
The updated seqbench_comparison.
Examples
design <- comparison_design(margin = 0.05, bounds = c(0, 1), boundary = "hoeffding")
state <- initialize_comparison(design)
set.seed(1)
losses <- data.frame(instance = 1:50, loss_a = runif(50, 0, 0.5),
loss_b = runif(50, 0.3, 0.8))
state <- update_comparison(state, losses)
stopping_decision(state)