Getting started with tidymatrix

Many datasets are a matrix with two tables attached: survey answers with information about the respondents and the questions, gene expression with gene and sample annotations, sales per shop and product. Kept as three separate objects, every filter or reordering has to be applied to all three at once, and it is easy to end up with a matrix that no longer lines up with its metadata.

tidymatrix keeps the three pieces together in one object and lets you manipulate it with the dplyr verbs you already know. Much like tidygraph does for networks, you activate() the part you want to work on (rows, columns or the matrix itself) and the rest follows along.

This vignette is a quick tour. Each section shows a few self-contained examples and links to a vignette that covers the topic in depth.

library(tidymatrix)
library(dplyr, warn.conflicts = FALSE)
#> Warning: package 'dplyr' was built under R version 4.5.2

The example data

All vignettes use big5, a simulated personality survey that ships with the package. 400 people answered 30 statements such as “I worry about things” on a scale from 1 (strongly disagree) to 5 (strongly agree). Each statement measures one of the Big Five personality traits, and two statements per trait are reverse-keyed: agreeing with “I keep in the background” means low extraversion.

The data come as three pieces: the answers,

big5_responses[1:5, 1:8]
#>      E1 E2 E3 E4 E5 E6 A1 A2
#> R001  4  3  5  3  3  4  3  5
#> R002  2  4  2  2  3  4  3  2
#> R003  2  2  3  4  4  3  3  3
#> R004  2  3  3  3  3  4  3  2
#> R005  1  2  2  1  5  5  5  5

what we know about the respondents,

head(big5_respondents)
#>   respondent_id age gender education     occupation country life_satisfaction
#> 1          R001  52 Female Secondary Service/Manual      EE                 8
#> 2          R002  32 Female    Master         Office      FI                 4
#> 3          R003  66   Male  Bachelor         Office      EE                 6
#> 4          R004  42   Male  Bachelor   Professional      EE                 9
#> 5          R005  48 Female Secondary         Office      FI                 8
#> 6          R006  62 Female     Basic Service/Manual      EE                 4
#>   completion_min
#> 1           19.8
#> 2            6.2
#> 3            6.3
#> 4            9.7
#> 5           14.2
#> 6           11.6

and what we know about the questions.

head(big5_items)
#>   item_id        trait reversed                             item_text position
#> 1      E1 Extraversion    FALSE     I feel comfortable around people.        1
#> 2      E2 Extraversion    FALSE                I start conversations.        6
#> 3      E3 Extraversion    FALSE I enjoy being part of a lively crowd.       11
#> 4      E4 Extraversion    FALSE                I make friends easily.       16
#> 5      E5 Extraversion     TRUE             I keep in the background.       21
#> 6      E6 Extraversion     TRUE    I have little to say to strangers.       26

tidymatrix() puts them together. Row metadata must have one row per matrix row, column metadata one row per matrix column.

tm <- tidymatrix(big5_responses, big5_respondents, big5_items)
tm
#> # A tidymatrix: 400 x 30 matrix
#> # Active: matrix
#> #
#> # Row data: 400 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Matrix preview:
#>      E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3 C4 C5 C6 N1 N2 N3 N4 N5 N6 O1
#> R001  4  3  5  3  3  4  3  5  3  4  4  3  1  2  1  2  4  5  3  3  3  2  3  2  2
#> R002  2  4  2  2  3  4  3  2  3  3  2  4  3  4  3  2  4  4  5  4  5  5  2  2  5
#> R003  2  2  3  4  4  3  3  3  4  5  4  3  2  3  1  3  4  4  4  4  3  2  4  2  1
#> R004  2  3  3  3  3  4  3  2  3  2  2  3  3  2  4  2  4  4  5  4  3  3  1  1  4
#> R005  1  2  2  1  5  5  5  5  5  5  2  1  2  5  4  3  2  2  4  2  3  4  3  2  2
#> R006  1  3  3  3  4  3  5  4  5  3  3  2  3  3  1  4  5  5  5  5  5  5  1  1  2
#>      O2 O3 O4 O5 O6
#> R001  4  1  3  2  3
#> R002  3  4  5  3  1
#> R003  4  1  2  5  4
#> R004  3  4  4  2  2
#> R005  3  2  3  4  3
#> R006  1  1  1  3  5

Working with rows and columns

activate(rows) makes dplyr verbs act on the respondents, activate(columns) on the questions. The matrix is subset and reordered to match automatically.

About 3% of the respondents clicked through the survey in a couple of minutes. Removing them is a filter on the rows, but shows up also on the dimensions of the matrix, as the rows are filtered:

tm_clean <- tm |>
  activate(rows) |>
  filter(completion_min > 3.5)

tm_clean |> 
  activate(matrix)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: matrix
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Matrix preview:
#>      E1 E2 E3 E4 E5 E6 A1 A2 A3 A4 A5 A6 C1 C2 C3 C4 C5 C6 N1 N2 N3 N4 N5 N6 O1
#> R001  4  3  5  3  3  4  3  5  3  4  4  3  1  2  1  2  4  5  3  3  3  2  3  2  2
#> R002  2  4  2  2  3  4  3  2  3  3  2  4  3  4  3  2  4  4  5  4  5  5  2  2  5
#> R003  2  2  3  4  4  3  3  3  4  5  4  3  2  3  1  3  4  4  4  4  3  2  4  2  1
#> R004  2  3  3  3  3  4  3  2  3  2  2  3  3  2  4  2  4  4  5  4  3  3  1  1  4
#> R005  1  2  2  1  5  5  5  5  5  5  2  1  2  5  4  3  2  2  4  2  3  4  3  2  2
#> R006  1  3  3  3  4  3  5  4  5  3  3  2  3  3  1  4  5  5  5  5  5  5  1  1  2
#>      O2 O3 O4 O5 O6
#> R001  4  1  3  2  3
#> R002  3  4  5  3  1
#> R003  4  1  2  5  4
#> R004  3  4  4  2  2
#> R005  3  2  3  4  3
#> R006  1  1  1  3  5

Keep only the neuroticism items, in the order they were asked:

tm_clean |>
  activate(columns) |>
  filter(trait == "Neuroticism") |>
  arrange(position)
#> # A tidymatrix: 388 x 6 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 6 rows x 5 columns
#> #
#> # Active data (columns):
#>   item_id       trait reversed                   item_text position
#> 1      N1 Neuroticism    FALSE  I get stressed out easily.        4
#> 2      N2 Neuroticism    FALSE       I worry about things.        9
#> 3      N3 Neuroticism    FALSE      My mood changes often.       14
#> 4      N4 Neuroticism    FALSE     I get irritated easily.       19
#> 5      N5 Neuroticism     TRUE I stay calm under pressure.       24
#> 6      N6 Neuroticism     TRUE         I seldom feel blue.       29

mutate(), select(), rename(), slice() and friends work the same way. Grouping and summarising aggregate the matrix as well as the metadata. Here we summarise the respondents by country: the metadata gets one row per country, and the matrix one row per country holding the average answer to each item.

by_country <- tm_clean |>
  activate(rows) |>
  group_by(country) |>
  summarise(n = n(), mean_age = mean(age))

by_country
#> # A tidymatrix: 6 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 6 rows x 3 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#>   country   n mean_age
#> 1      DE  11 43.90909
#> 2      EE 159 42.59119
#> 3      FI  71 41.36620
#> 4      LT  46 44.36957
#> 5      LV  50 39.30000
#> 6      SE  51 41.50980

by_country |>
  activate(matrix) |>
  transform_matrix(round, digits = 2)
#> # A tidymatrix: 6 x 30 matrix
#> # Active: matrix
#> #
#> # Row data: 6 rows x 3 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Matrix preview:
#>        E1   E2   E3   E4   E5   E6   A1   A2   A3   A4   A5   A6   C1   C2   C3
#> [1,] 2.55 3.09 2.82 3.45 3.09 2.82 3.64 3.91 3.55 3.18 3.18 2.00 3.91 3.00 3.45
#> [2,] 2.97 3.16 3.34 3.30 2.97 3.28 3.17 3.64 3.21 3.74 3.36 2.43 3.23 2.65 2.64
#> [3,] 3.15 3.04 3.30 2.99 3.07 3.37 3.37 3.69 3.49 3.70 3.14 2.35 3.39 2.72 2.69
#> [4,] 2.83 3.00 3.80 3.28 2.91 3.41 3.35 3.50 3.26 3.78 3.33 2.30 3.59 2.87 2.59
#> [5,] 2.70 3.02 3.34 3.32 2.60 3.44 2.94 3.72 3.30 3.56 3.54 2.42 3.58 2.76 2.64
#> [6,] 3.06 3.08 3.35 3.37 2.94 3.12 3.02 3.33 3.25 3.73 3.25 2.41 3.39 2.57 2.88
#>        C4   C5   C6   N1   N2   N3   N4   N5   N6   O1   O2   O3   O4   O5   O6
#> [1,] 3.82 3.00 3.27 2.91 3.09 2.91 2.82 3.18 2.73 3.18 3.36 2.55 3.09 3.36 2.82
#> [2,] 3.06 3.57 3.20 3.17 2.89 3.07 3.09 3.10 2.35 3.21 3.57 2.84 3.47 3.01 3.12
#> [3,] 3.20 3.59 3.07 3.20 3.15 3.24 3.34 2.99 2.20 3.06 3.38 2.56 3.38 3.07 3.25
#> [4,] 3.09 3.65 3.09 2.87 2.78 2.83 3.02 3.24 2.65 3.54 4.20 3.02 3.61 2.67 2.87
#> [5,] 3.24 3.48 3.30 3.16 3.04 3.02 3.14 2.98 2.24 3.00 3.36 2.64 3.52 2.92 3.02
#> [6,] 3.27 3.49 3.16 3.35 2.78 3.20 3.27 2.94 2.29 3.22 3.57 3.04 3.49 3.12 3.10

More: vignette("dplyr-verbs", package = "tidymatrix") — Working with rows and columns.

Joins

Joins add information from another table to the active metadata. The package includes a small country table:

big5_countries
#>   country country_name region   language population_m
#> 1      EE      Estonia Baltic   Estonian         1.37
#> 2      FI      Finland Nordic    Finnish         5.60
#> 3      LV       Latvia Baltic    Latvian         1.87
#> 4      LT    Lithuania Baltic Lithuanian         2.89
#> 5      SE       Sweden Nordic    Swedish        10.55
#> 6      NO       Norway Nordic  Norwegian         5.55
tm_clean |>
  activate(rows) |>
  left_join(big5_countries, by = "country") |>
  select(respondent_id, country, country_name, region)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 4 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#>   respondent_id country country_name region
#> 1          R001      EE      Estonia Baltic
#> 2          R002      FI      Finland Nordic
#> 3          R003      EE      Estonia Baltic
#> 4          R004      EE      Estonia Baltic
#> 5          R005      FI      Finland Nordic
#> 6          R006      EE      Estonia Baltic

Germany is missing from the country table, so German respondents get NA. anti_join() finds them, and inner_join() would drop them together with their rows of the matrix.

tm_clean |>
  activate(rows) |>
  anti_join(big5_countries, by = "country") |>
  pull(country) |>
  table()
#> 
#> DE 
#> 11

More: Joins.

Matrix operations

Questionnaires usually word some statements in the opposite direction, so that people who agree with everything do not automatically get a high score. These are called reverse-keyed items. In big5, “I start conversations” measures extraversion directly, but “I keep in the background” is reverse-keyed: strongly agreeing with it (5) means low extraversion. Before items can be combined, the answers to reverse-keyed items are flipped (1 ↔︎ 5, 2 ↔︎ 4, 3 stays), so that a high value always means a high trait level. The reversed column of the item metadata says which items need flipping.

transform_matrix() applies a function to the whole matrix, to each row or to each column, depending on what is active. With rows active the function receives one respondent’s answers, and extra arguments can refer to the column metadata — here the reversed flag of each item:

tm_scored <- tm_clean |>
  activate(rows) |>
  transform_matrix(\(x, flip) ifelse(flip, 6L - x, x), flip = reversed)

In tm_scored, averaging a respondent’s answers within a trait now gives a meaningful trait score. Scaling, centering, clipping, transposing and summary statistics are also available:

tm_scored |>
  activate(columns) |>
  scale() |>                 # z-score each item
  activate(matrix) |>
  clip_values(min = -2, max = 2) |>
  activate(rows) |>
  add_stats(mean, sd)        # per-respondent mean and sd into row metadata
#> # A tidymatrix: 388 x 30 matrix
#> # Active: rows
#> #
#> # Row data: 388 rows x 10 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (rows):
#>   respondent_id age gender education     occupation country life_satisfaction
#> 1          R001  52 Female Secondary Service/Manual      EE                 8
#> 2          R002  32 Female    Master         Office      FI                 4
#> 3          R003  66   Male  Bachelor         Office      EE                 6
#> 4          R004  42   Male  Bachelor   Professional      EE                 9
#> 5          R005  48 Female Secondary         Office      FI                 8
#> 6          R006  62 Female     Basic Service/Manual      EE                 4
#>   completion_min        mean        sd
#> 1           19.8 -0.26387972 0.8062588
#> 2            6.2  0.17412754 1.0271633
#> 3            6.3 -0.41261943 0.8113876
#> 4            9.7  0.05753829 0.8781575
#> 5           14.2  0.01124931 1.1579789
#> 6           11.6 -0.05557806 1.2445335

More: Matrix operations.

PCA, clustering and MDS

Analysis functions store their results where they belong: coordinates and cluster labels become metadata columns, and the full result object is kept for later. With columns active, each item becomes a point:

tm_items <- tm_scored |>
  activate(columns) |>
  compute_prcomp(n_components = 2) |>
  compute_hclust(k = 5, method = "ward.D2") |>
  compute_mds(k = 2)

tm_items |>
  activate(columns) |>
  select(item_id, trait, column_pca_PC1, column_pca_PC2, column_hclust_cluster)
#> # A tidymatrix: 388 x 30 matrix
#> # Active: columns
#> #
#> # Row data: 388 rows x 8 columns
#> # Column data: 30 rows x 5 columns
#> #
#> # Active data (columns):
#>    item_id        trait column_pca_PC1 column_pca_PC2 column_hclust_cluster
#> E1      E1 Extraversion     -2.0811342      -7.498614                     1
#> E2      E2 Extraversion     -0.4268466      -6.344608                     1
#> E3      E3 Extraversion      0.7957248      -3.490065                     1
#> E4      E4 Extraversion     -0.8497496      -4.218130                     1
#> E5      E5 Extraversion     -2.1985266      -6.838678                     1
#> E6      E6 Extraversion     -3.2929751     -11.292714                     1

list_analyses(tm_items)
#> [1] "column_pca"    "column_hclust" "column_mds"

The clustering recovers the five traits from the answers alone:

tm_items |>
  activate(columns) |>
  as_tibble() |>
  count(trait, column_hclust_cluster)
#> # A tibble: 5 × 3
#>   trait             column_hclust_cluster     n
#>   <chr>             <fct>                 <int>
#> 1 Agreeableness     2                         6
#> 2 Conscientiousness 3                         6
#> 3 Extraversion      1                         6
#> 4 Neuroticism       4                         6
#> 5 Openness          5                         6

More: PCA and clustering and MDS, t-SNE and UMAP.

Statistics for every row or column

compute_ttest(), compute_correlation(), compute_lm() and others fit one test per item (or per respondent) and return a tidy table. Do women and men answer differently?

gender_tests <- tm_scored |>
  activate(rows) |>
  filter(gender %in% c("Female", "Male")) |>
  activate(columns) |>
  compute_ttest(group_col = "gender", control = "Male", treatment = "Female")

gender_tests |>
  filter(p.adj < 0.05) |>
  select(item_id, trait, item_text, p.value, p.adj)
#> # A tibble: 9 × 5
#>   item_id trait         item_text                                p.value   p.adj
#>   <chr>   <chr>         <chr>                                      <dbl>   <dbl>
#> 1 A3      Agreeableness I trust what people tell me.             1.03e-3 5.14e-3
#> 2 A4      Agreeableness I am quick to forgive.                   1.45e-2 4.85e-2
#> 3 A6      Agreeableness I am not really interested in others' p… 9.50e-3 3.56e-2
#> 4 N1      Neuroticism   I get stressed out easily.               5.35e-5 5.35e-4
#> 5 N2      Neuroticism   I worry about things.                    4.24e-3 1.82e-2
#> 6 N3      Neuroticism   My mood changes often.                   1.54e-6 3.03e-5
#> 7 N4      Neuroticism   I get irritated easily.                  7.17e-5 5.38e-4
#> 8 N5      Neuroticism   I stay calm under pressure.              2.02e-6 3.03e-5
#> 9 N6      Neuroticism   I seldom feel blue.                      4.28e-4 2.57e-3

The differences concentrate on agreeableness and neuroticism, as in real personality data. compute_across() lets you run any function of your own in the same way.

More: Row- and column-wise statistics.

Plotting and exporting

The metadata is an ordinary data frame, so plotting the item PCA is one as_tibble() away:

library(ggplot2)
#> Warning: package 'ggplot2' was built under R version 4.5.2

tm_items |>
  activate(columns) |>
  as_tibble() |>
  ggplot(aes(column_pca_PC1, column_pca_PC2, colour = trait, label = item_id)) +
  geom_text() +
  labs(x = "PC1", y = "PC2", colour = NULL) +
  theme_minimal()

to_long() turns the whole object into one row per matrix cell, with all the metadata attached, ready for ggplot2 or for writing to a file:

tm_scored |>
  to_long() |>
  select(respondent_id, gender, item_id, trait, value) |>
  head()
#> # A tibble: 6 × 5
#>   respondent_id gender item_id trait        value
#>   <chr>         <chr>  <chr>   <chr>        <int>
#> 1 R001          Female E1      Extraversion     4
#> 2 R002          Female E1      Extraversion     2
#> 3 R003          Male   E1      Extraversion     2
#> 4 R004          Male   E1      Extraversion     2
#> 5 R005          Female E1      Extraversion     1
#> 6 R006          Female E1      Extraversion     1

plot_pheatmap() draws a heatmap with the metadata as annotations and reuses stored clusterings:

tm_scored |>
  activate(columns) |>
  scale() |>
  compute_hclust(method = "ward.D2") |>
  activate(rows) |>
  compute_hclust(method = "ward.D2") |>
  plot_pheatmap(
    col_names = "item_id",
    row_annotation = c("gender", "age"),
    col_annotation = "trait",
    show_rownames = FALSE
  )

More: Plotting and exporting.

Where next

Vignette Covers
Working with rows and columns filter(), mutate(), arrange(), slice(), grouping and summarising
Joins Adding external annotations; all six join types
Matrix operations Transformations, scaling, clipping, transposing, statistics
PCA and clustering compute_prcomp(), compute_hclust(), compute_kmeans(), stored analyses
MDS, t-SNE and UMAP Non-linear and distance-based embeddings
Row- and column-wise statistics t-tests, correlations, linear models, ANOVA, custom functions
Plotting and exporting ggplot2, heatmaps, to_long(), getting data out