| Title: | Score Comparability for Translated and Adapted Exams |
| Version: | 0.1.0 |
| Description: | An adaptation-comparability workflow for small, lower-scoring and unbalanced language groups where standard differential item functioning (DIF) tools (Magis, Beland, Tuerlinckx and De Boeck, 2010, <doi:10.3758/BRM.42.3.847>) break down. Calibrates the Rasch model in each language group, links the groups robustly through the densest cluster of items rather than assuming DIF cancels out (on anchor selection see Kopf, Zeileis and Strobl, 2015, <doi:10.1177/0013164414529792>), detects small-sample DIF with an empirical-Bayes spike-and-slab model and local false discovery rates (Efron, 2004, <doi:10.1198/016214504000000089>), quantifies whether item-level DIF accumulates into different pass rates, explains DIF by item features to give translators actionable guidance, and drafts a comparability report. |
| License: | MIT + file LICENSE |
| URL: | https://github.com/edidatasolutions/transDIF, https://edidatasolutions.github.io/transDIF/ |
| BugReports: | https://github.com/edidatasolutions/transDIF/issues |
| Encoding: | UTF-8 |
| Depends: | R (≥ 4.1) |
| Imports: | stats |
| Suggests: | knitr, markdown |
| VignetteBuilder: | knitr |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-28 00:45:20 UTC; User |
| Author: | Daniel Edi |
| Maintainer: | Daniel Edi <danieledi2026@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-08 09:20:02 UTC |
Calibrate each language group separately
Description
Fits the Rasch model by marginal maximum likelihood within each group, each with its own mean-zero ability scale. The between-language difference 'd = b_focal - b_ref' therefore equals item DIF plus a common shift 'c = -(focal mean ability on the reference scale)'. Separating the two is the linking problem [td_dif()] solves.
Usage
td_calibrate(responses, group, ref = "ref", focal = "focal")
Arguments
responses |
0/1 matrix (persons x items), column names = item ids. |
group |
Group label per person. |
ref, focal |
Labels of the reference (source-language) and focal (translated) groups. |
Value
An 'td_calibration': data frame 'items' ('item', 'b_ref', 'se_ref', 'b_focal', 'se_focal', 'd', 'se_d') plus group sizes, ability SDs and ‘loc_var' (variance of the two scales’ locations). 'se_ref'/'se_focal' are absolute SEs; 'se_d' uses relative SEs, excluding the location uncertainty that is common to all items (it belongs to the linking shift).
Examples
sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
cal <- td_calibrate(sim$responses, sim$group)
head(cal$items)
Robust linking, anchor selection and small-sample DIF
Description
The between-language difference of item 'i' is modeled as 'd_i ~ N(c + delta_i, se_i^2)', where 'c' is the common shift (the ability difference) that linking must recover and ‘delta_i' is the item’s DIF.
Usage
td_dif(
calibration,
fdr = 0.1,
tau0 = 0.05,
anchor_max = 0.2,
link = c("mode", "mixture"),
ambiguity = 0.8
)
Arguments
calibration |
A 'td_calibration'. |
fdr |
Target Bayesian false discovery rate for flagging. |
tau0 |
SD of DIF among "DIF-free" items: the scale of DIF considered negligible (logits). |
anchor_max |
Items with posterior DIF probability below this are reported as anchors. |
link |
'"mode"' (default) or '"mixture"'. |
ambiguity |
Competing-mode density ratio above which the linking is reported as ambiguous. |
Details
**Linking.** By default 'c' is the precision-weighted *mode* of the 'd_i': the center of the densest cluster of items, found by maximizing 'sum_i phi((d_i - c) / s_i) / s_i' with 's_i^2 = se_i^2 + tau0^2'. It assumes that DIF-free items form the largest cluster, not that DIF cancels out as mean linking does. In known-truth benchmarks across balanced, directional and heavy DIF, it had the lowest overall RMSE of the estimators compared (including mean linking, iterative purification, median, Tukey biweight, least trimmed squares and the joint mixture below). Its SE is a sandwich estimate plus the scales' location variance. If a second cluster of items is nearly as dense, the linking is flagged as ambiguous and the competing shift is reported.
**DIF.** Given 'c', DIF effects follow a spike-and-slab mixture: DIF-free items (share 'pi0 > 0.5') have 'delta_i ~ N(0, tau0^2)', and DIF items have 'delta_i ~ N(m1, tau1^2)', where 'm1' allows directional DIF. Each item gets a posterior DIF probability, a shrunken DIF estimate and a local false discovery rate. 'link = "mixture"' instead estimates 'c' jointly in the mixture (the approach of version 0.1.0; kept for comparison).
Value
A 'td_dif' object: '$items' ('item', 'd', 'se_d', 'p_dif', 'lfdr', 'dif_mean', 'dif_sd' (posterior mean/SD of DIF), 'flag', 'anchor'), '$link' ('c', 'c_se', 'pi0', 'm1', 'tau0', 'tau1', 'alt_c', 'alt_ratio'), baseline linking constants 'c_mean' (all items as anchors) and 'c_purified', and 'weakly_identified' (TRUE, with a warning, when a second item cluster is nearly as dense as the chosen one).
Examples
sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
dif
# true linking shift is -focal_mean = 0.5
c(estimate = dif$link[["c"]], mean_linking = dif$c_mean)
Which item features predict DIF?
Description
Random-effects meta-regression of each item's DIF estimate ('d - c') on item features (e.g. idioms, cultural referents, measurement units, vocabulary load), weighting by '1 / (se_d^2 + tau^2)'. The between-item variance 'tau^2' not explained by the features is estimated by the method of moments. Coefficients are logits of DIF per unit of the feature. That is guidance a translation team can act on.
Usage
td_features(dif, features)
Arguments
dif |
An 'td_dif'. |
features |
Data frame with 'item' and numeric feature columns. |
Value
Data frame of coefficients ('term', 'estimate', 'se', 'z', 'p_value') with attribute 'tau' (residual DIF SD). Features with no variation across items are dropped with a warning.
Examples
sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
td_features(dif, sim$features)
Does item-level DIF add up to different pass rates?
Description
Differential test functioning for the translated form. With difficulties on the reference scale ('b_ref') and DIF effects 'delta', the translated form has difficulties 'b_ref + delta'. For the translated-language population (ability N(-c, sigma_focal^2) on the reference scale) the function computes the pass rate at raw cut 'cut' on the translated form versus on a DIF-free form, using exact score distributions. It also computes the expected raw-score shift for an examinee exactly at the cut. Intervals propagate both the linking uncertainty (the focal group's mean ability is '-c') and the posterior uncertainty of each item's DIF.
Usage
td_impact(dif, cut, n_draws = 200, level = 0.9, seed = NULL)
Arguments
dif |
An 'td_dif'. |
cut |
Raw-score passing standard. |
n_draws |
Posterior draws. |
level |
Interval level. |
seed |
Optional seed. |
Value
A data frame with estimate and interval for 'pass_rate_fair', 'pass_rate_translated', 'pass_rate_change' and 'score_shift_at_cut'.
Examples
sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
td_impact(dif, cut = 12, n_draws = 50, seed = 1)
Mantel-Haenszel DIF with purification (conventional baseline)
Description
Mantel-Haenszel DIF with purification (conventional baseline)
Usage
td_mh(responses, group, ref = "ref", focal = "focal", fdr = 0.1, purify = TRUE)
Arguments
responses |
0/1 matrix. |
group |
Group label per person. |
ref, focal |
Group labels. |
fdr |
Benjamini-Hochberg level for flagging. |
purify |
Re-run once matching on the total over non-flagged items. |
Value
Data frame: 'item', 'alpha_mh', 'delta_mh' (ETS delta scale), 'p_value', 'flag'.
Examples
sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
mh <- td_mh(sim$responses, sim$group)
table(flagged = mh$flag, true_dif = sim$truth$dif_item)
Draft a comparability report
Description
Assembles a plain-language Markdown report of the linking, DIF, aggregate impact and feature analyses, organized around the kinds of evidence the Standards for Educational and Psychological Testing (2014; fairness chapter) and the ITC Guidelines for Translating and Adapting Tests (2nd ed., 2017) ask for. It is a draft for a psychometrician to review, not a finished validity argument.
Usage
td_report(
dif,
impact = NULL,
features = NULL,
languages = c("source", "target"),
file = NULL
)
Arguments
dif |
An 'td_dif'. |
impact |
Optional output of [td_impact()]. |
features |
Optional output of [td_features()]. |
languages |
Names of the source and target languages. |
file |
Optional path to write the Markdown to. |
Value
The report as a character string (invisibly if written to file).
Examples
sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
rep <- td_report(dif, td_impact(dif, cut = 12, n_draws = 50, seed = 1),
td_features(dif, sim$features), c("English", "French"))
cat(rep)
# To save: td_report(dif, file = file.path(tempdir(), "comparability.md"))
Simulate a source-language and a translated administration with known DIF
Description
Items carry binary adaptation features (idiom, cultural referent, measurement units, high vocabulary load). In the translated form an item's difficulty shifts by 'sum(feature_effects * features)' plus small noise, so DIF is directional (translations mostly harder) and unbalanced, which is the case where mean-based linking fails. The focal (translated) group is small and lower-scoring on average.
Usage
td_simulate(
n_ref = 2000,
n_focal = 150,
n_items = 60,
focal_mean = -0.5,
focal_sd = 1,
feature_prev = c(idiom = 0.1, cultural = 0.1, units = 0.08, vocabulary = 0.12),
feature_effects = c(idiom = 0.6, cultural = 0.5, units = -0.4, vocabulary = 0.35),
dif_noise = 0.1,
seed = NULL
)
Arguments
n_ref, n_focal |
Group sizes. |
n_items |
Test length. |
focal_mean, focal_sd |
Focal-group ability (reference is N(0, 1)). |
feature_prev |
Prevalence of each feature. |
feature_effects |
DIF (logits) contributed by each feature. |
dif_noise |
SD of feature-unrelated DIF on flagged items. |
seed |
Optional seed. |
Value
An 'td_sim': '$responses' (0/1 matrix, persons x items), '$group' ('"ref"'/'"focal"'), '$features' (item data frame), '$truth' ('b_ref', 'dif', 'focal_mean', 'focal_sd').
Examples
sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
table(dif_item = sim$truth$dif_item)
head(sim$features)