---
title: "Would this candidate have passed with different raters?"
output: markdown::html_format
vignette: >
  %\VignetteIndexEntry{Would this candidate have passed with different raters?}
  %\VignetteEngine{knitr::knitr}
  %\VignetteEncoding{UTF-8}
---

```{r, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```

Boards that run oral exams, OSCEs or essay-scored certifications must defend
individual pass/fail decisions. Rater severity is well studied; what it *did*
to decisions usually is not. decisionfacets answers that question directly.

## A simulated administration

Every candidate is scored on four tasks by two raters drawn from a pool of
twelve whose severities differ. Because the data are simulated, the true
abilities and severities are known.

```{r}
library(decisionfacets)
sim <- df_simulate(n_persons = 400, n_items = 4, n_raters = 12,
                   raters_per_person = 2, severity_sd = 0.6, seed = 2026)
head(sim$data)
round(sim$par$lambda, 2)   # true rater severities (positive = harsher)
```

## Fit the many-facet Rasch model

`df_fit()` uses 'TAM' when it is installed and a built-in joint maximum
likelihood estimator otherwise. We use the built-in one here so the vignette
has no dependencies.

```{r}
fit <- df_fit(sim$data, engine = "jmle")
df_rater_effects(fit)
```

## The decision rule is the point

The same substantive standard (an average rating of 2 per cell) can be applied
to raw totals or to severity-adjusted measures. Rater severity reaches the
decision only in the first case.

```{r}
raw <- df_cut(2 * 4 * 2, "raw_total")
fa  <- df_cut(2, "fair_average")
```

## Counterfactual pass probabilities

For each candidate, `df_counterfactual()` gives the probability of passing a
re-rating by the observed panel, by an average-severity panel and by a random
panel from the pool. `advantage` is how far the assigned panel pushed the
candidate toward the outcome they received.

```{r}
cf_raw <- df_counterfactual(fit, raw)
cf_raw
summary(cf_raw)
summary(df_counterfactual(fit, fa))
```

Under raw totals, a sizable share of candidates are rater-dependent; under the
fair average, essentially none are.

## Where do wrong decisions come from?

```{r}
df_attribute(cf_raw)
```

The decomposition separates error that no rater could remove (measurement)
from error added by the particular raters assigned, and from the expected cost
of random assignment.

## Checking against the truth

With the true parameters the same analysis gives the known answer, so the
estimates can be compared with it:

```{r}
truth <- df_counterfactual(sim, raw)
cor(truth$advantage, cf_raw$advantage)
table(truth = truth$rater_dependent, estimated = cf_raw$rater_dependent)
```
