Many datasets are a matrix with two tables attached: survey answers with information about the respondents and the questions, gene expression with gene and sample annotations, sales per shop and product. Kept as three separate objects, every filter or reordering has to be applied to all three at once, and it is easy to end up with a matrix that no longer lines up with its metadata.
tidymatrix keeps the three pieces together in one object
and lets you manipulate it with the dplyr verbs you already know. Much
like tidygraph does for networks, you
activate() the part you want to work on (rows, columns or
the matrix itself) and the rest follows along.
This vignette is a quick tour. Each section shows a few self-contained examples and links to a vignette that covers the topic in depth.
library(tidymatrix)
library(dplyr, warn.conflicts = FALSE)
#> Warning: package 'dplyr' was built under R version 4.5.2All vignettes use big5, a simulated personality survey
that ships with the package. 400 people answered 30 statements such as
“I worry about things” on a scale from 1 (strongly disagree) to
5 (strongly agree). Each statement measures one of the Big Five
personality traits, and two statements per trait are
reverse-keyed: agreeing with “I keep in the
background” means low extraversion.
The data come as three pieces: the answers,
big5_responses[1:5, 1:8]
#> E1 E2 E3 E4 E5 E6 A1 A2
#> R001 4 3 5 3 3 4 3 5
#> R002 2 4 2 2 3 4 3 2
#> R003 2 2 3 4 4 3 3 3
#> R004 2 3 3 3 3 4 3 2
#> R005 1 2 2 1 5 5 5 5what we know about the respondents,
head(big5_respondents)
#> respondent_id age gender education occupation country life_satisfaction
#> 1 R001 52 Female Secondary Service/Manual EE 8
#> 2 R002 32 Female Master Office FI 4
#> 3 R003 66 Male Bachelor Office EE 6
#> 4 R004 42 Male Bachelor Professional EE 9
#> 5 R005 48 Female Secondary Office FI 8
#> 6 R006 62 Female Basic Service/Manual EE 4
#> completion_min
#> 1 19.8
#> 2 6.2
#> 3 6.3
#> 4 9.7
#> 5 14.2
#> 6 11.6and what we know about the questions.
head(big5_items)
#> item_id trait reversed item_text position
#> 1 E1 Extraversion FALSE I feel comfortable around people. 1
#> 2 E2 Extraversion FALSE I start conversations. 6
#> 3 E3 Extraversion FALSE I enjoy being part of a lively crowd. 11
#> 4 E4 Extraversion FALSE I make friends easily. 16
#> 5 E5 Extraversion TRUE I keep in the background. 21
#> 6 E6 Extraversion TRUE I have little to say to strangers. 26tidymatrix() puts them together. Row metadata must have
one row per matrix row, column metadata one row per matrix column.
tm <- tidymatrix(big5_responses, big5_respondents, big5_items)
tm
#> # A tidymatrix: 400 x 30 matrix
#> # Active: matrix
#> #
#> # Row data: 400 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Matrix preview:
#> E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3 C4 C5 C6 N1 N2 N3 N4 N5 N6 O1
#> R001 4 3 5 3 3 4 3 5 3 4 4 3 1 2 1 2 4 5 3 3 3 2 3 2 2
#> R002 2 4 2 2 3 4 3 2 3 3 2 4 3 4 3 2 4 4 5 4 5 5 2 2 5
#> R003 2 2 3 4 4 3 3 3 4 5 4 3 2 3 1 3 4 4 4 4 3 2 4 2 1
#> R004 2 3 3 3 3 4 3 2 3 2 2 3 3 2 4 2 4 4 5 4 3 3 1 1 4
#> R005 1 2 2 1 5 5 5 5 5 5 2 1 2 5 4 3 2 2 4 2 3 4 3 2 2
#> R006 1 3 3 3 4 3 5 4 5 3 3 2 3 3 1 4 5 5 5 5 5 5 1 1 2
#> O2 O3 O4 O5 O6
#> R001 4 1 3 2 3
#> R002 3 4 5 3 1
#> R003 4 1 2 5 4
#> R004 3 4 4 2 2
#> R005 3 2 3 4 3
#> R006 1 1 1 3 5activate(rows) makes dplyr verbs act on the respondents,
activate(columns) on the questions. The matrix is subset
and reordered to match automatically.
About 3% of the respondents clicked through the survey in a couple of minutes. Removing them is a filter on the rows, but shows up also on the dimensions of the matrix, as the rows are filtered:
tm_clean <- tm |>
activate(rows) |>
filter(completion_min > 3.5)
tm_clean |>
activate(matrix)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: matrix
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Matrix preview:
#> E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3 C4 C5 C6 N1 N2 N3 N4 N5 N6 O1
#> R001 4 3 5 3 3 4 3 5 3 4 4 3 1 2 1 2 4 5 3 3 3 2 3 2 2
#> R002 2 4 2 2 3 4 3 2 3 3 2 4 3 4 3 2 4 4 5 4 5 5 2 2 5
#> R003 2 2 3 4 4 3 3 3 4 5 4 3 2 3 1 3 4 4 4 4 3 2 4 2 1
#> R004 2 3 3 3 3 4 3 2 3 2 2 3 3 2 4 2 4 4 5 4 3 3 1 1 4
#> R005 1 2 2 1 5 5 5 5 5 5 2 1 2 5 4 3 2 2 4 2 3 4 3 2 2
#> R006 1 3 3 3 4 3 5 4 5 3 3 2 3 3 1 4 5 5 5 5 5 5 1 1 2
#> O2 O3 O4 O5 O6
#> R001 4 1 3 2 3
#> R002 3 4 5 3 1
#> R003 4 1 2 5 4
#> R004 3 4 4 2 2
#> R005 3 2 3 4 3
#> R006 1 1 1 3 5Keep only the neuroticism items, in the order they were asked:
tm_clean |>
activate(columns) |>
filter(trait == "Neuroticism") |>
arrange(position)
#> # A tidymatrix: 388 x 6 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 6 rows x 5 columns
#> #
#> # Active data (columns):
#> item_id trait reversed item_text position
#> 1 N1 Neuroticism FALSE I get stressed out easily. 4
#> 2 N2 Neuroticism FALSE I worry about things. 9
#> 3 N3 Neuroticism FALSE My mood changes often. 14
#> 4 N4 Neuroticism FALSE I get irritated easily. 19
#> 5 N5 Neuroticism TRUE I stay calm under pressure. 24
#> 6 N6 Neuroticism TRUE I seldom feel blue. 29mutate(), select(), rename(),
slice() and friends work the same way. Grouping and
summarising aggregate the matrix as well as the metadata. Here we
summarise the respondents by country: the metadata gets one row per
country, and the matrix one row per country holding the average answer
to each item.
by_country <- tm_clean |>
activate(rows) |>
group_by(country) |>
summarise(n = n(), mean_age = mean(age))
by_country
#> # A tidymatrix: 6 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 6 rows x 3 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#> country n mean_age
#> 1 DE 11 43.90909
#> 2 EE 159 42.59119
#> 3 FI 71 41.36620
#> 4 LT 46 44.36957
#> 5 LV 50 39.30000
#> 6 SE 51 41.50980
by_country |>
activate(matrix) |>
transform_matrix(round, digits = 2)
#> # A tidymatrix: 6 x 30 matrix
#> # Active: matrix
#> #
#> # Row data: 6 rows x 3 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Matrix preview:
#> E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3
#> [1,] 2.55 3.09 2.82 3.45 3.09 2.82 3.64 3.91 3.55 3.18 3.18 2.00 3.91 3.00 3.45
#> [2,] 2.97 3.16 3.34 3.30 2.97 3.28 3.17 3.64 3.21 3.74 3.36 2.43 3.23 2.65 2.64
#> [3,] 3.15 3.04 3.30 2.99 3.07 3.37 3.37 3.69 3.49 3.70 3.14 2.35 3.39 2.72 2.69
#> [4,] 2.83 3.00 3.80 3.28 2.91 3.41 3.35 3.50 3.26 3.78 3.33 2.30 3.59 2.87 2.59
#> [5,] 2.70 3.02 3.34 3.32 2.60 3.44 2.94 3.72 3.30 3.56 3.54 2.42 3.58 2.76 2.64
#> [6,] 3.06 3.08 3.35 3.37 2.94 3.12 3.02 3.33 3.25 3.73 3.25 2.41 3.39 2.57 2.88
#> C4 C5 C6 N1 N2 N3 N4 N5 N6 O1 O2 O3 O4 O5 O6
#> [1,] 3.82 3.00 3.27 2.91 3.09 2.91 2.82 3.18 2.73 3.18 3.36 2.55 3.09 3.36 2.82
#> [2,] 3.06 3.57 3.20 3.17 2.89 3.07 3.09 3.10 2.35 3.21 3.57 2.84 3.47 3.01 3.12
#> [3,] 3.20 3.59 3.07 3.20 3.15 3.24 3.34 2.99 2.20 3.06 3.38 2.56 3.38 3.07 3.25
#> [4,] 3.09 3.65 3.09 2.87 2.78 2.83 3.02 3.24 2.65 3.54 4.20 3.02 3.61 2.67 2.87
#> [5,] 3.24 3.48 3.30 3.16 3.04 3.02 3.14 2.98 2.24 3.00 3.36 2.64 3.52 2.92 3.02
#> [6,] 3.27 3.49 3.16 3.35 2.78 3.20 3.27 2.94 2.29 3.22 3.57 3.04 3.49 3.12 3.10More:
vignette("dplyr-verbs", package = "tidymatrix") — Working with rows and columns.
Joins add information from another table to the active metadata. The package includes a small country table:
big5_countries
#> country country_name region language population_m
#> 1 EE Estonia Baltic Estonian 1.37
#> 2 FI Finland Nordic Finnish 5.60
#> 3 LV Latvia Baltic Latvian 1.87
#> 4 LT Lithuania Baltic Lithuanian 2.89
#> 5 SE Sweden Nordic Swedish 10.55
#> 6 NO Norway Nordic Norwegian 5.55tm_clean |>
activate(rows) |>
left_join(big5_countries, by = "country") |>
select(respondent_id, country, country_name, region)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 4 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#> respondent_id country country_name region
#> 1 R001 EE Estonia Baltic
#> 2 R002 FI Finland Nordic
#> 3 R003 EE Estonia Baltic
#> 4 R004 EE Estonia Baltic
#> 5 R005 FI Finland Nordic
#> 6 R006 EE Estonia BalticGermany is missing from the country table, so German respondents get
NA. anti_join() finds them, and
inner_join() would drop them together with their rows of
the matrix.
tm_clean |>
activate(rows) |>
anti_join(big5_countries, by = "country") |>
pull(country) |>
table()
#>
#> DE
#> 11More: Joins.
Questionnaires usually word some statements in the opposite
direction, so that people who agree with everything do not automatically
get a high score. These are called reverse-keyed items. In
big5, “I start conversations” measures
extraversion directly, but “I keep in the background” is
reverse-keyed: strongly agreeing with it (5) means low
extraversion. Before items can be combined, the answers to reverse-keyed
items are flipped (1 ↔︎ 5, 2 ↔︎ 4, 3 stays), so that a high value always
means a high trait level. The reversed column of the item
metadata says which items need flipping.
transform_matrix() applies a function to the whole
matrix, to each row or to each column, depending on what is active. With
rows active the function receives one respondent’s answers, and extra
arguments can refer to the column metadata — here the
reversed flag of each item:
tm_scored <- tm_clean |>
activate(rows) |>
transform_matrix(\(x, flip) ifelse(flip, 6L - x, x), flip = reversed)In tm_scored, averaging a respondent’s answers within a
trait now gives a meaningful trait score. Scaling, centering, clipping,
transposing and summary statistics are also available:
tm_scored |>
activate(columns) |>
scale() |> # z-score each item
activate(matrix) |>
clip_values(min = -2, max = 2) |>
activate(rows) |>
add_stats(mean, sd) # per-respondent mean and sd into row metadata
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#> respondent_id age gender education occupation country life_satisfaction
#> 1 R001 52 Female Secondary Service/Manual EE 8
#> 2 R002 32 Female Master Office FI 4
#> 3 R003 66 Male Bachelor Office EE 6
#> 4 R004 42 Male Bachelor Professional EE 9
#> 5 R005 48 Female Secondary Office FI 8
#> 6 R006 62 Female Basic Service/Manual EE 4
#> completion_min mean sd
#> 1 19.8 -0.26387972 0.8062588
#> 2 6.2 0.17412754 1.0271633
#> 3 6.3 -0.41261943 0.8113876
#> 4 9.7 0.05753829 0.8781575
#> 5 14.2 0.01124931 1.1579789
#> 6 11.6 -0.05557806 1.2445335More: Matrix operations.
Analysis functions store their results where they belong: coordinates and cluster labels become metadata columns, and the full result object is kept for later. With columns active, each item becomes a point:
tm_items <- tm_scored |>
activate(columns) |>
compute_prcomp(n_components = 2) |>
compute_hclust(k = 5, method = "ward.D2") |>
compute_mds(k = 2)
tm_items |>
activate(columns) |>
select(item_id, trait, column_pca_PC1, column_pca_PC2, column_hclust_cluster)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (columns):
#> item_id trait column_pca_PC1 column_pca_PC2 column_hclust_cluster
#> E1 E1 Extraversion -2.0811342 -7.498614 1
#> E2 E2 Extraversion -0.4268466 -6.344608 1
#> E3 E3 Extraversion 0.7957248 -3.490065 1
#> E4 E4 Extraversion -0.8497496 -4.218130 1
#> E5 E5 Extraversion -2.1985266 -6.838678 1
#> E6 E6 Extraversion -3.2929751 -11.292714 1
list_analyses(tm_items)
#> [1] "column_pca" "column_hclust" "column_mds"The clustering recovers the five traits from the answers alone:
tm_items |>
activate(columns) |>
as_tibble() |>
count(trait, column_hclust_cluster)
#> # A tibble: 5 × 3
#> trait column_hclust_cluster n
#> <chr> <fct> <int>
#> 1 Agreeableness 2 6
#> 2 Conscientiousness 3 6
#> 3 Extraversion 1 6
#> 4 Neuroticism 4 6
#> 5 Openness 5 6More: PCA and clustering and MDS, t-SNE and UMAP.
compute_ttest(), compute_correlation(),
compute_lm() and others fit one test per item (or per
respondent) and return a tidy table. Do women and men answer
differently?
gender_tests <- tm_scored |>
activate(rows) |>
filter(gender %in% c("Female", "Male")) |>
activate(columns) |>
compute_ttest(group_col = "gender", control = "Male", treatment = "Female")
gender_tests |>
filter(p.adj < 0.05) |>
select(item_id, trait, item_text, p.value, p.adj)
#> # A tibble: 9 × 5
#> item_id trait item_text p.value p.adj
#> <chr> <chr> <chr> <dbl> <dbl>
#> 1 A3 Agreeableness I trust what people tell me. 1.03e-3 5.14e-3
#> 2 A4 Agreeableness I am quick to forgive. 1.45e-2 4.85e-2
#> 3 A6 Agreeableness I am not really interested in others' p… 9.50e-3 3.56e-2
#> 4 N1 Neuroticism I get stressed out easily. 5.35e-5 5.35e-4
#> 5 N2 Neuroticism I worry about things. 4.24e-3 1.82e-2
#> 6 N3 Neuroticism My mood changes often. 1.54e-6 3.03e-5
#> 7 N4 Neuroticism I get irritated easily. 7.17e-5 5.38e-4
#> 8 N5 Neuroticism I stay calm under pressure. 2.02e-6 3.03e-5
#> 9 N6 Neuroticism I seldom feel blue. 4.28e-4 2.57e-3The differences concentrate on agreeableness and neuroticism, as in
real personality data. compute_across() lets you run any
function of your own in the same way.
The metadata is an ordinary data frame, so plotting the item PCA is
one as_tibble() away:
library(ggplot2)
#> Warning: package 'ggplot2' was built under R version 4.5.2
tm_items |>
activate(columns) |>
as_tibble() |>
ggplot(aes(column_pca_PC1, column_pca_PC2, colour = trait, label = item_id)) +
geom_text() +
labs(x = "PC1", y = "PC2", colour = NULL) +
theme_minimal()to_long() turns the whole object into one row per matrix
cell, with all the metadata attached, ready for ggplot2 or for writing
to a file:
tm_scored |>
to_long() |>
select(respondent_id, gender, item_id, trait, value) |>
head()
#> # A tibble: 6 × 5
#> respondent_id gender item_id trait value
#> <chr> <chr> <chr> <chr> <int>
#> 1 R001 Female E1 Extraversion 4
#> 2 R002 Female E1 Extraversion 2
#> 3 R003 Male E1 Extraversion 2
#> 4 R004 Male E1 Extraversion 2
#> 5 R005 Female E1 Extraversion 1
#> 6 R006 Female E1 Extraversion 1plot_pheatmap() draws a heatmap with the metadata as
annotations and reuses stored clusterings:
tm_scored |>
activate(columns) |>
scale() |>
compute_hclust(method = "ward.D2") |>
activate(rows) |>
compute_hclust(method = "ward.D2") |>
plot_pheatmap(
col_names = "item_id",
row_annotation = c("gender", "age"),
col_annotation = "trait",
show_rownames = FALSE
)More: Plotting and exporting.
| Vignette | Covers |
|---|---|
| Working with rows and columns | filter(), mutate(),
arrange(), slice(), grouping and
summarising |
| Joins | Adding external annotations; all six join types |
| Matrix operations | Transformations, scaling, clipping, transposing, statistics |
| PCA and clustering | compute_prcomp(), compute_hclust(),
compute_kmeans(), stored analyses |
| MDS, t-SNE and UMAP | Non-linear and distance-based embeddings |
| Row- and column-wise statistics | t-tests, correlations, linear models, ANOVA, custom functions |
| Plotting and exporting | ggplot2, heatmaps, to_long(), getting data out |