| Title: | A Unified Tidy Interface to R's Machine Learning Ecosystem |
| Version: | 0.6.0 |
| Description: | Provides a unified tidyverse-compatible interface to R's machine learning ecosystem - from data ingestion to model publishing. The tl_read() family reads data from files ('CSV', 'Excel', 'Parquet', 'JSON'), databases ('SQLite', 'PostgreSQL', 'MySQL', 'BigQuery'), and cloud sources ('S3', 'GitHub', 'Kaggle'). The tl_model() function wraps established implementations from 'glmnet', 'randomForest', 'xgboost', 'e1071', 'rpart', 'gbm', 'nnet', 'cluster', 'dbscan', and others with consistent function signatures and tidy tibble output. Results flow into unified 'ggplot2'-based visualization and optional formatted 'gt' tables via the tl_table() family. The underlying algorithms are unchanged; 'tidylearn' simply makes them easier to use together. Access raw model objects via the $fit slot for a supervised method, or $fit$model for an unsupervised one. Methods include random forests Breiman (2001) <doi:10.1023/A:1010933404324>, LASSO regression Tibshirani (1996) <doi:10.1111/j.2517-6161.1996.tb02080.x>, elastic net Zou and Hastie (2005) <doi:10.1111/j.1467-9868.2005.00503.x>, support vector machines Cortes and Vapnik (1995) <doi:10.1007/BF00994018>, and gradient boosting Friedman (2001) <doi:10.1214/aos/1013203451>. |
| License: | MIT + file LICENSE |
| Encoding: | UTF-8 |
| RoxygenNote: | 7.3.3 |
| Depends: | R (≥ 4.1.0) |
| Imports: | dplyr (≥ 1.0.0), ggplot2 (≥ 3.4.0), tibble (≥ 3.0.0), tidyr (≥ 1.0.0), purrr (≥ 0.3.0), rlang (≥ 1.0.0), magrittr, parallel, stats, e1071, gbm, glmnet, nnet, randomForest, rpart, rsample, ROCR, yardstick, cluster (≥ 2.1.0), dbscan (≥ 1.1.0), MASS, smacof (≥ 2.1.0) |
| Suggests: | arules, arulesViz, bigrquery, car, DBI, DiagrammeR, DT, GGally, ggforce, gridExtra, gt, httr2 (≥ 1.0.0), jsonlite, keras, knitr, lmtest, moments, nanoparquet, NeuralNetTools, paws.storage, readr, RhpcBLASctl, readxl, RMariaDB, rmarkdown, RPostgres, rpart.plot, RSQLite, shiny, shinydashboard, tensorflow, testthat (≥ 3.1.7), withr, xgboost |
| Config/testthat/edition: | 3 |
| URL: | https://tidylearn.sheetsolved.com, https://github.com/ces0491/tidylearn |
| BugReports: | https://github.com/ces0491/tidylearn/issues |
| VignetteBuilder: | knitr |
| Collate: | 'tidylearn-package.R' 'utils.R' 'read.R' 'read-backends.R' 'core.R' 'preprocessing.R' 'supervised-classification.R' 'supervised-regression.R' 'supervised-regularization.R' 'supervised-trees.R' 'supervised-svm.R' 'supervised-neural-networks.R' 'supervised-deep-learning.R' 'supervised-xgboost.R' 'unsupervised-distance.R' 'unsupervised-pca.R' 'unsupervised-mds.R' 'unsupervised-clustering.R' 'unsupervised-hclust.R' 'unsupervised-dbscan.R' 'unsupervised-market-basket.R' 'unsupervised-validation.R' 'integration.R' 'pipeline.R' 'model-selection.R' 'compute-detection.R' 'compute-advisor.R' 'compute-routing.R' 'cloud-consent.R' 'cloud-endpoint.R' 'cloud-cost.R' 'cloud-serialize.R' 'tuning.R' 'interactions.R' 'diagnostics.R' 'metrics.R' 'coefficients.R' 'visualization.R' 'tables.R' 'workflows.R' |
| NeedsCompilation: | no |
| Packaged: | 2026-10-08 15:50:50 UTC; CesaireTobias |
| Author: | Cesaire Tobias [aut, cre] |
| Maintainer: | Cesaire Tobias <cesaire@sheetsolved.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-08 18:20:02 UTC |
tidylearn: A Unified Tidy Interface to R's Machine Learning Ecosystem
Description
Provides a unified tidyverse-compatible interface to R's machine learning ecosystem - from data ingestion to model publishing. The tl_read() family reads data from files ('CSV', 'Excel', 'Parquet', 'JSON'), databases ('SQLite', 'PostgreSQL', 'MySQL', 'BigQuery'), and cloud sources ('S3', 'GitHub', 'Kaggle'). The tl_model() function wraps established implementations from 'glmnet', 'randomForest', 'xgboost', 'e1071', 'rpart', 'gbm', 'nnet', 'cluster', 'dbscan', and others with consistent function signatures and tidy tibble output. Results flow into unified 'ggplot2'-based visualization and optional formatted 'gt' tables via the tl_table() family. The underlying algorithms are unchanged; 'tidylearn' simply makes them easier to use together. Access raw model objects via the $fit slot for a supervised method, or $fit$model for an unsupervised one. Methods include random forests Breiman (2001) doi:10.1023/A:1010933404324, LASSO regression Tibshirani (1996) doi:10.1111/j.2517-6161.1996.tb02080.x, elastic net Zou and Hastie (2005) doi:10.1111/j.1467-9868.2005.00503.x, support vector machines Cortes and Vapnik (1995) doi:10.1007/BF00994018, and gradient boosting Friedman (2001) doi:10.1214/aos/1013203451.
Details
tidylearn wraps established R machine learning packages behind one consistent interface. The main entry points are:
tl_readRead data from files, databases and cloud sources into a tidy tibble.
tl_modelFit any supported supervised or unsupervised method.
tl_evaluateScore a fitted model.
tl_tableRender results as formatted gt tables.
tl_auto_mlSearch across methods automatically.
Every fitted model keeps the underlying package's own object in its
$fit slot, so package-specific functionality remains available.
See vignette("getting-started", package = "tidylearn") for a
walkthrough.
Author(s)
Maintainer: Cesaire Tobias cesaire@sheetsolved.com
See Also
Useful links:
Report bugs at https://github.com/ces0491/tidylearn/issues
Pipe operator
Description
See magrittr::%>% for details.
Usage
lhs %>% rhs
lhs %>% rhs
Arguments
lhs |
A value or the magrittr placeholder. |
rhs |
A function call using the magrittr semantics. |
Value
The result of applying rhs to lhs.
Augment Data with DBSCAN Cluster Assignments
Description
Augment Data with DBSCAN Cluster Assignments
Usage
augment_dbscan(dbscan_obj, data)
Arguments
dbscan_obj |
A tidy_dbscan object |
data |
Original data frame |
Value
A tibble containing the original data with additional columns
cluster (factor), is_noise (logical), and is_core
(logical).
Examples
db <- tidy_dbscan(iris[, 1:4], eps = 0.5, minPts = 5)
augmented <- augment_dbscan(db, iris)
Augment Data with Hierarchical Cluster Assignments
Description
Add cluster assignments to original data
Usage
augment_hclust(hclust_obj, data, k = NULL, h = NULL)
Arguments
hclust_obj |
A tidy_hclust object |
data |
Original data frame |
k |
Number of clusters (optional) |
h |
Height at which to cut (optional) |
Value
A tibble containing the original data with an additional
cluster integer column indicating cluster assignments.
Examples
hc <- tidy_hclust(USArrests, method = "ward.D2")
augmented <- augment_hclust(hc, USArrests, k = 3)
Augment Data with K-Means Cluster Assignments
Description
Augment Data with K-Means Cluster Assignments
Usage
augment_kmeans(kmeans_obj, data)
Arguments
kmeans_obj |
A tidy_kmeans object |
data |
Original data frame |
Value
A tibble containing the original data with an additional
cluster factor column indicating cluster assignments.
Examples
km <- tidy_kmeans(iris[, 1:4], k = 3)
augmented <- augment_kmeans(km, iris)
Augment Data with PAM Cluster Assignments
Description
Augment Data with PAM Cluster Assignments
Usage
augment_pam(pam_obj, data)
Arguments
pam_obj |
A tidy_pam object |
data |
Original data frame |
Value
A tibble containing the original data with an additional
cluster factor column indicating cluster assignments.
Examples
pm <- tidy_pam(iris[, 1:4], k = 3)
augmented <- augment_pam(pm, iris)
Augment Original Data with PCA Scores
Description
Add PC scores to the original dataset
Usage
augment_pca(pca_obj, data, n_components = NULL)
Arguments
pca_obj |
A tidy_pca object |
data |
Original data frame |
n_components |
Number of PCs to add (default: all) |
Value
A tibble containing the original data with additional columns
for each principal component score (named PC1, PC2, etc.).
Examples
pca <- tidy_pca(USArrests)
augmented <- augment_pca(pca, USArrests, n_components = 2)
Calculate Cluster Validation Metrics
Description
Comprehensive validation metrics for a clustering result
Usage
calc_validation_metrics(clusters, data = NULL, dist_mat = NULL)
Arguments
clusters |
Vector of cluster assignments: numeric, factor or
character. A label of 0 marks noise, as |
data |
Original data frame (for WSS calculation). WSS is taken over its numeric columns, so it needs at least one. |
dist_mat |
Distance matrix (for silhouette) |
Value
A single-row tibble with columns k, min_size,
max_size, avg_size, n_noise, and optionally
avg_silhouette, min_silhouette (if dist_mat
provided; NA for a single cluster), and total_wss (if
data provided).
Examples
km <- kmeans(iris[, 1:4], centers = 3, nstart = 25)
d <- dist(iris[, 1:4])
metrics <- calc_validation_metrics(km$cluster, iris[, 1:4], d)
Calculate Within-Cluster Sum of Squares for Different k
Description
Used for elbow method to determine optimal k
Usage
calc_wss(data, max_k = 10, nstart = 25)
Arguments
data |
A data frame or tibble |
max_k |
Maximum number of clusters to test (default: 10) |
nstart |
Number of random starts for each k (default: 25) |
Value
A tibble with columns k (number of clusters) and
tot_withinss (total within-cluster sum of squares).
Examples
wss <- calc_wss(iris[, 1:4], max_k = 6)
plot(wss$k, wss$tot_withinss, type = "b")
Compare Multiple Clustering Results
Description
Compare Multiple Clustering Results
Usage
compare_clusterings(cluster_list, data, dist_mat = NULL)
Arguments
cluster_list |
Named list of cluster assignment vectors. An entry
without a name is reported as |
data |
Original data |
dist_mat |
Distance matrix |
Value
A tibble with one row per clustering method and columns for each
validation metric (see calc_validation_metrics), plus a
method column identifying the clustering.
Examples
km3 <- kmeans(iris[, 1:4], 3, nstart = 25)$cluster
km4 <- kmeans(iris[, 1:4], 4, nstart = 25)$cluster
compare_clusterings(list(k3 = km3, k4 = km4), iris[, 1:4])
Compare Distance Methods
Description
Compute distances using multiple methods for comparison
Usage
compare_distances(data, methods = c("euclidean", "manhattan", "maximum"))
Arguments
data |
A data frame or tibble |
methods |
Character vector of methods to compare |
Value
A named list of dist objects, one per method.
Examples
dists <- compare_distances(
iris[, 1:4], methods = c("euclidean", "manhattan")
)
Create Summary Dashboard
Description
Generate a multi-panel summary of clustering results
Usage
create_cluster_dashboard(
data,
cluster_col = "cluster",
validation_metrics = NULL
)
Arguments
data |
Data frame with cluster assignments |
cluster_col |
Cluster column name |
validation_metrics |
Optional tibble of validation metrics |
Value
Invisibly returns a named list of the
ggplot objects drawn: clusters, the
scatter plot, when the data has two numeric columns besides
cluster_col; sizes; and metrics, when
validation_metrics is given. The combined plot grid is drawn as
a side effect via grid.arrange.
Examples
df <- iris[, 1:4]
df$cluster <- kmeans(df, 3)$cluster
create_cluster_dashboard(df)
Explore DBSCAN Parameters
Description
Test multiple eps and minPts combinations
Usage
explore_dbscan_params(data, eps_values, minPts_values)
Arguments
data |
A data frame or matrix |
eps_values |
Vector of eps values to test |
minPts_values |
Vector of minPts values to test |
Value
A tibble with columns eps, minPts, n_clusters,
n_noise, and prop_noise for each parameter combination.
Examples
params <- explore_dbscan_params(iris[, 1:4],
eps_values = c(0.3, 0.5, 0.8), minPts_values = c(3, 5))
Filter Rules by Item
Description
Subset rules containing specific items
Usage
filter_rules_by_item(rules_obj, item, where = "both")
Arguments
rules_obj |
A tidy_apriori object, an arules rules object, or a
tibble of rules from |
item |
Character; one item name, matched against whole items, so "coffee" does not match "instant coffee" |
where |
Character; "lhs", "rhs", or "both" (default: "both") |
Value
A tibble of rules containing the specified item in the
requested position.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
res <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
filter_rules_by_item(res, "whole milk", where = "rhs")
}
Find Related Items
Description
Find items frequently purchased with a given item
Usage
find_related_items(rules_obj, item, min_lift = 1.5, top_n = 10)
Arguments
rules_obj |
A tidy_apriori object, an arules rules object, or a
tibble of rules from |
item |
Character; one item name to find associations for, matched against whole items |
min_lift |
Minimum lift threshold (default: 1.5) |
top_n |
Number of top associations to return (default: 10) |
Value
A tibble of rules involving the specified item, filtered by
min_lift and sorted by lift in descending order.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
res <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
find_related_items(res, "whole milk", min_lift = 1.5)
}
Get PCA Loadings in Wide Format
Description
Get PCA Loadings in Wide Format
Usage
get_pca_loadings(pca_obj, n_components = NULL)
Arguments
pca_obj |
A |
n_components |
Number of components to include (default: all) |
Value
A tibble with one row per variable and one column per principal component, containing the loading values.
Examples
pca <- tidy_pca(USArrests)
get_pca_loadings(pca, n_components = 2)
Get Variance Explained Summary
Description
Get Variance Explained Summary
Usage
get_pca_variance(pca_obj)
Arguments
pca_obj |
A |
Value
A tibble with columns component, sdev,
variance, prop_variance, and cum_variance.
Examples
pca <- tidy_pca(USArrests)
get_pca_variance(pca)
# The same accessor works on a tl_model() PCA fit
get_pca_variance(tl_model(USArrests, method = "pca"))
Inspect Association Rules
Description
View rules sorted by various quality measures
Usage
inspect_rules(rules_obj, by = "lift", n = 10, decreasing = TRUE)
Arguments
rules_obj |
A tidy_apriori object, an arules rules or itemsets object, or a tibble of rules |
by |
Sort by: "support", "confidence", "lift" (default), "count". Itemsets have no lift, so for them the default sorts by support. |
n |
Number of rules to display (default: 10) |
decreasing |
If TRUE (default), the |
Value
A tibble of the n rules ranked highest (or, with
decreasing = FALSE, lowest) by the quality measure by.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
res <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
inspect_rules(res, by = "lift", n = 5)
}
Find Optimal Number of Clusters
Description
Use multiple methods to suggest optimal k
Usage
optimal_clusters(data, max_k = 10, methods = c("silhouette", "gap", "wss"))
Arguments
data |
A data frame or tibble |
max_k |
Maximum k to test (default: 10) |
methods |
Vector of methods: "silhouette", "gap", "wss" (default: all) |
Value
A list of class "optimal_k_results" containing one or more of:
wss: tibble from
calc_wss(if "wss" method used)silhouette: tibble from
tidy_silhouette_analysis(if "silhouette" method used)gap: a
tidy_gapobject fromtidy_gap_stat(if "gap" method used)
Examples
opt <- optimal_clusters(iris[, 1:4], max_k = 6, methods = "wss")
Determine Optimal Number of Clusters for Hierarchical Clustering
Description
Use silhouette or gap statistic to find optimal k
Usage
optimal_hclust_k(hclust_obj, method = "silhouette", max_k = 10)
Arguments
hclust_obj |
A tidy_hclust object |
method |
Character; "silhouette" (default) or "gap". The gap
statistic resamples the observations' numeric columns, so it refuses a
tree built from a dist object, and one built with
|
max_k |
Maximum number of clusters to test (default: 10) |
Value
A list containing:
optimal_k: the recommended number of clusters
method: the evaluation method used
values: numeric vector of evaluation scores (for silhouette)
k_range: integer vector of k values tested (for silhouette)
If method = "gap", returns a tidy_gap object instead.
Examples
hc <- tidy_hclust(USArrests, method = "ward.D2")
opt <- optimal_hclust_k(hc, method = "silhouette", max_k = 6)
Plot EDA results
Description
Plot EDA results
Usage
## S3 method for class 'tidylearn_eda'
plot(x, ...)
Arguments
x |
A tidylearn_eda object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly. Called for its
side effect of plotting a PCA scatter plot coloured by cluster.
Examples
eda <- tl_explore(iris, response = "Species")
plot(eda)
Plot method for tidylearn models
Description
Plot method for tidylearn models
Usage
## S3 method for class 'tidylearn_model'
plot(x, type = "auto", ...)
Arguments
x |
A tidylearn model object |
type |
Plot type (default: "auto") |
... |
Additional arguments passed to plotting functions |
Value
A ggplot object. The specific plot depends
on the model paradigm and type argument.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
plot(model, type = "actual_predicted")
Create Cluster Comparison Plot
Description
Compare multiple clustering results side-by-side
Usage
plot_cluster_comparison(data, cluster_cols, x_col, y_col)
Arguments
data |
Data frame with multiple cluster columns |
cluster_cols |
Vector of cluster column names |
x_col |
X-axis variable |
y_col |
Y-axis variable |
Value
The return value of grid.arrange, a
gtable drawn as a side effect.
Examples
df <- iris[, 1:4]
df$km3 <- kmeans(df, 3)$cluster
df$km4 <- kmeans(df, 4)$cluster
plot_cluster_comparison(df, c("km3", "km4"), "Sepal.Length", "Sepal.Width")
Plot Cluster Size Distribution
Description
Create bar plot of cluster sizes
Usage
plot_cluster_sizes(clusters, title = "Cluster Size Distribution")
Arguments
clusters |
Vector of cluster assignments |
title |
Plot title (default: "Cluster Size Distribution") |
Value
A ggplot object.
Examples
clusters <- kmeans(iris[, 1:4], 3)$cluster
plot_cluster_sizes(clusters)
Plot Clusters in 2D Space
Description
Visualize clustering results using first two dimensions or specified dimensions
Usage
plot_clusters(
data,
cluster_col = "cluster",
x_col = NULL,
y_col = NULL,
centers = NULL,
title = "Cluster Plot",
color_noise_black = TRUE
)
Arguments
data |
A data frame with cluster assignments |
cluster_col |
Name of cluster column (default: "cluster") |
x_col |
X-axis variable (if NULL, uses the first numeric column
other than |
y_col |
Y-axis variable (if NULL, uses the second numeric column
other than |
centers |
Optional data frame of cluster centers |
title |
Plot title |
color_noise_black |
If TRUE, color noise points (cluster 0) black |
Value
A ggplot object.
Examples
km <- tidy_kmeans(iris[, 1:4], k = 3)
clustered <- augment_kmeans(km, iris[, 1:4])
plot_clusters(clustered)
Plot Dendrogram with Cluster Highlights
Description
Enhanced dendrogram with colored cluster rectangles
Usage
plot_dendrogram(
hclust_obj,
k = NULL,
title = "Hierarchical Clustering Dendrogram"
)
Arguments
hclust_obj |
Hierarchical clustering object: an |
k |
Number of clusters to highlight |
title |
Plot title |
Value
Invisibly returns the hclust object. The
dendrogram is drawn as a side effect.
Examples
hc <- hclust(dist(iris[, 1:4]))
plot_dendrogram(hc, k = 3)
Create Distance Heatmap
Description
Visualize distance matrix as heatmap
Usage
plot_distance_heatmap(
dist_mat,
cluster_order = NULL,
title = "Distance Heatmap"
)
Arguments
dist_mat |
Distance matrix (dist object) |
cluster_order |
Optional vector to reorder observations by cluster |
title |
Plot title |
Value
A ggplot object.
Examples
d <- dist(iris[1:20, 1:4])
plot_distance_heatmap(d)
Create Elbow Plot for K-Means
Description
Plot total within-cluster sum of squares vs number of clusters
Usage
plot_elbow(wss_data, add_line = FALSE, suggested_k = NULL)
Arguments
wss_data |
A tibble with columns k and tot_withinss (from calc_wss) |
add_line |
Add vertical line at suggested optimal k? (default: FALSE) |
suggested_k |
If add_line=TRUE, which k to highlight |
Value
A ggplot object.
Examples
wss <- data.frame(k = 2:6, tot_withinss = c(150, 90, 60, 50, 45))
plot_elbow(wss)
Plot Gap Statistic
Description
Plot Gap Statistic
Usage
plot_gap_stat(gap_obj, show_methods = FALSE)
Arguments
gap_obj |
A tidy_gap object |
show_methods |
Logical; show all three k selection methods? (default: FALSE) |
Value
A ggplot object.
Examples
gap <- tidy_gap_stat(iris[, 1:4], max_k = 6, B = 10)
plot_gap_stat(gap)
Plot k-NN Distance Plot
Description
Visualize k-NN distances to help choose eps
Usage
plot_knn_dist(data, k = 4, add_suggestion = TRUE, percentile = 0.95)
Arguments
data |
A data frame, matrix, or tidy_knn_dist result |
k |
If data is a data frame, k for k-NN (default: 4) |
add_suggestion |
Add suggested eps line? (default: TRUE) |
percentile |
Percentile for suggestion (default: 0.95) |
Value
A ggplot object.
Examples
plot_knn_dist(iris[, 1:4], k = 5)
Plot MDS Configuration
Description
Visualize MDS results
Usage
plot_mds(mds_obj, color_by = NULL, label_points = TRUE, dim_x = 1, dim_y = 2)
Arguments
mds_obj |
A tidy_mds object |
color_by |
Optional grouping to colour points by: either a column name present in the MDS configuration, or a vector as long as the data. |
label_points |
Logical; add point labels? (default: TRUE) |
dim_x |
Which dimension for x-axis (default: 1) |
dim_y |
Which dimension for y-axis (default: 2) |
Value
A ggplot object.
Examples
mds <- tidy_mds(USArrests, method = "classical")
plot_mds(mds)
Plot Silhouette Analysis
Description
Plot Silhouette Analysis
Usage
plot_silhouette(sil_obj)
Arguments
sil_obj |
A tidy_silhouette object or tibble from tidy_silhouette_analysis |
Value
A ggplot object.
Examples
km <- kmeans(iris[, 1:4], centers = 3, nstart = 25)
d <- dist(iris[, 1:4])
sil <- tidy_silhouette(km$cluster, d)
plot_silhouette(sil)
Plot Variance Explained (PCA)
Description
Create combined scree plot showing individual and cumulative variance
Usage
plot_variance_explained(variance_tbl, threshold = 0.8)
Arguments
variance_tbl |
Variance tibble from tidy_pca |
threshold |
Horizontal line for variance threshold (default: 0.8 for 80%) |
Value
A ggplot object.
Examples
model <- tl_model(iris[, 1:4], method = "pca")
plot_variance_explained(model$fit$variance_explained)
Predict using a tidylearn model
Description
Unified prediction interface for both supervised and unsupervised models
Usage
## S3 method for class 'tidylearn_model'
predict(object, new_data = NULL, type = "response", ...)
Arguments
object |
A tidylearn model object |
new_data |
A data frame containing the new data. If NULL, uses training data. |
type |
Type of prediction, for supervised models only:
|
... |
Additional arguments |
Value
For supervised models, a tibble with a
.pred column; with type = "prob", one column per class
instead. It has one row per row of new_data, in order, and
zero rows for zero-row new_data. For unsupervised models, the
method's natural output: an .obs_id column, the row names of
the data predicted on, plus component scores for "pca" and
"mds" (as many as the model keeps), or plus a cluster
column for the clustering methods.
Unsupervised models differ in whether they can handle new data.
"pca" projects it and "kmeans" assigns it to the
nearest centre; "pam", "clara", "dbscan",
"mds" and "hclust" have no out-of-sample projection
and error if new_data is supplied. For hierarchical
clustering, cut the tree with tidy_cutree() instead.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
predict(model)
predict(model, new_data = mtcars[1:5, ])
Predict from stratified models
Description
Predict from stratified models
Usage
## S3 method for class 'tidylearn_stratified'
predict(object, new_data = NULL, ...)
Arguments
object |
A tidylearn_stratified model object |
new_data |
New data for predictions. NULL predicts the training rows from their stored cluster assignments, which works for every clustering method; new rows can be assigned by k-means only. |
... |
Additional arguments passed to each cluster's model, such as
|
Value
A tibble of the columns each cluster's model
returns for the requested type – .pred by default, one
column per class for type = "prob" – and a .cluster
column with cluster assignments. Rows of a single-class cluster are
predicted as that class, with probability 1.
Examples
models <- tl_stratified_models(mtcars, mpg ~ .,
cluster_method = "kmeans", k = 2, supervised_method = "linear")
preds <- predict(models)
Predict with transfer learning model
Description
Predict with transfer learning model
Usage
## S3 method for class 'tidylearn_transfer'
predict(object, new_data, ...)
Arguments
object |
A tidylearn_transfer model object |
new_data |
New data for predictions |
... |
Additional arguments |
Value
A tibble with a .pred column containing
predictions.
Examples
model <- tl_transfer_learning(iris, Species ~ .,
pretrain_method = "pca", supervised_method = "tree")
preds <- predict(model, iris[1:5, ])
Print Method for tidy_apriori
Description
Print Method for tidy_apriori
Usage
## S3 method for class 'tidy_apriori'
print(x, ...)
Arguments
x |
A tidy_apriori object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
res <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
print(res)
}
Print Method for tidy_dbscan
Description
Print Method for tidy_dbscan
Usage
## S3 method for class 'tidy_dbscan'
print(x, ...)
Arguments
x |
A tidy_dbscan object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
db <- tidy_dbscan(iris[, 1:4], eps = 0.5, minPts = 5)
print(db)
Print Method for tidy_gap
Description
Print Method for tidy_gap
Usage
## S3 method for class 'tidy_gap'
print(x, ...)
Arguments
x |
A tidy_gap object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
gap <- tidy_gap_stat(iris[, 1:4], max_k = 6, B = 10)
print(gap)
Print Method for tidy_hclust
Description
Print Method for tidy_hclust
Usage
## S3 method for class 'tidy_hclust'
print(x, ...)
Arguments
x |
A tidy_hclust object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
hc <- tidy_hclust(USArrests, method = "ward.D2")
print(hc)
Print Method for tidy_kmeans
Description
Print Method for tidy_kmeans
Usage
## S3 method for class 'tidy_kmeans'
print(x, ...)
Arguments
x |
A tidy_kmeans object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
km <- tidy_kmeans(iris[, 1:4], k = 3)
print(km)
Print Method for tidy_mds
Description
Print Method for tidy_mds
Usage
## S3 method for class 'tidy_mds'
print(x, ...)
Arguments
x |
A tidy_mds object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
mds <- tidy_mds(USArrests, method = "classical")
print(mds)
Print Method for tidy_pam
Description
Print Method for tidy_pam
Usage
## S3 method for class 'tidy_pam'
print(x, ...)
Arguments
x |
A tidy_pam object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
pm <- tidy_pam(iris[, 1:4], k = 3)
print(pm)
Print Method for tidy_pca
Description
Print Method for tidy_pca
Usage
## S3 method for class 'tidy_pca'
print(x, ...)
Arguments
x |
A tidy_pca object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
pca <- tidy_pca(USArrests)
print(pca)
Print Method for tidy_silhouette
Description
Print Method for tidy_silhouette
Usage
## S3 method for class 'tidy_silhouette'
print(x, ...)
Arguments
x |
A tidy_silhouette object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
km <- kmeans(iris[, 1:4], centers = 3, nstart = 25)
d <- dist(iris[, 1:4])
sil <- tidy_silhouette(km$cluster, d)
print(sil)
Print auto ML results
Description
Print auto ML results
Usage
## S3 method for class 'tidylearn_automl'
print(x, ...)
Arguments
x |
A tidylearn_automl object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
result <- tl_auto_ml(iris, Species ~ .,
time_budget = 10,
use_reduction = FALSE,
use_clustering = FALSE,
cv_folds = 2)
# The leaderboard, the winner and the metric it was ranked on
print(result)
Print method for tidylearn_compute_advice objects
Description
Print method for tidylearn_compute_advice objects
Usage
## S3 method for class 'tidylearn_compute_advice'
print(x, ...)
Arguments
x |
A |
... |
Unused. |
Value
The input x, invisibly.
Examples
# Runtime, peak memory and cost per tier, before committing to a fit
advice <- tl_compute_advisor("forest", data = iris, formula = Species ~ .)
print(advice)
Print a tidylearn_data object
Description
Print a tidylearn_data object
Usage
## S3 method for class 'tidylearn_data'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments passed to the tibble print method. |
Value
The input object x, returned invisibly.
Examples
f <- tempfile(fileext = ".csv")
write.csv(iris, f, row.names = FALSE)
d <- tl_read(f)
print(d)
unlink(f)
Print EDA results
Description
Print EDA results
Usage
## S3 method for class 'tidylearn_eda'
print(x, ...)
Arguments
x |
A tidylearn_eda object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
eda <- tl_explore(iris, response = "Species")
print(eda)
Print method for tidylearn_gpu_check objects
Description
Print method for tidylearn_gpu_check objects
Usage
## S3 method for class 'tidylearn_gpu_check'
print(x, ...)
Arguments
x |
A |
... |
Unused. |
Value
The input x, invisibly.
Examples
# Reports no GPU rather than failing on a machine without one
print(tl_check_gpu())
Print method for tidylearn models
Description
Print method for tidylearn models
Usage
## S3 method for class 'tidylearn_model'
print(x, ...)
Arguments
x |
A tidylearn model object |
... |
Additional arguments (ignored) |
Value
The input object x, returned invisibly.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
print(model)
Print a tidylearn pipeline
Description
Print a tidylearn pipeline
Usage
## S3 method for class 'tidylearn_pipeline'
print(x, ...)
Arguments
x |
A tidylearn pipeline object |
... |
Additional arguments (not used) |
Value
The input pipeline object x, returned invisibly.
Examples
pipe <- tl_pipeline(iris, Species ~ .)
print(pipe)
Generate Product Recommendations
Description
Get product recommendations based on basket contents
Usage
recommend_products(rules_obj, basket, top_n = 5, min_confidence = 0.5)
Arguments
rules_obj |
A tidy_apriori object, an arules rules object, or a
tibble of rules from |
basket |
Character vector of items in current basket |
top_n |
Number of recommendations to return (default: 5) |
min_confidence |
Minimum confidence threshold (default: 0.5) |
Value
A tibble with columns rhs (recommended item),
confidence, lift, and support, sorted by lift in
descending order. A rule is used when the basket holds its whole
left-hand side and none of its right-hand side, and each product is
listed once, from its highest-lift rule.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
res <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
# The basket has to cover the whole left-hand side of a rule, so a
# basket of very common items usually matches nothing above the
# confidence floor
recommend_products(res, basket = c("flour", "baking powder"))
}
Standardize Data
Description
Center and/or scale numeric variables
Usage
standardize_data(data, center = TRUE, scale = TRUE)
Arguments
data |
A data frame or tibble |
center |
Logical; center variables? (default: TRUE) |
scale |
Logical; scale variables to unit variance? (default: TRUE) |
Value
A tibble with numeric variables centered and/or scaled as specified;
non-numeric columns are returned unchanged. A grouped tibble is
standardised within each group, as dplyr::mutate() works on it;
a rowwise tibble is standardised over its whole columns, since a single
value has no spread. Grouping and rowwise identifier columns are left
as they are.
Examples
std <- standardize_data(iris[, 1:4])
Suggest eps Parameter for DBSCAN
Description
Use k-NN distance plot to suggest eps value
Usage
suggest_eps(data, minPts = 5, method = "percentile", percentile = 0.95)
Arguments
data |
A data frame or matrix |
minPts |
The |
method |
Method to suggest eps: "percentile" (default), "knee" |
percentile |
If method="percentile", which percentile to use (default: 0.95) |
Value
A list containing:
eps: suggested epsilon value
knn_distances: full tibble of k-NN distances
method: method used
Examples
eps_info <- suggest_eps(iris, minPts = 5)
eps_info$eps
Summarize Association Rules
Description
Get summary statistics about rules
Usage
summarize_rules(rules_obj)
Arguments
rules_obj |
A tidy_apriori object, an arules rules object, or a rules tibble |
Value
A list with n_rules and summary statistics (min,
max, mean, median) for support,
confidence, and lift.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
res <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
summarize_rules(res)
}
Summary method for tidylearn models
Description
Summary method for tidylearn models
Usage
## S3 method for class 'tidylearn_model'
summary(object, ...)
Arguments
object |
A tidylearn model object |
... |
Additional arguments (ignored) |
Value
The input object, returned invisibly. Called for its
side effect of printing model summary and training performance.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
summary(model)
Summarize a tidylearn pipeline
Description
Summarize a tidylearn pipeline
Usage
## S3 method for class 'tidylearn_pipeline'
summary(object, ...)
Arguments
object |
A tidylearn pipeline object |
... |
Additional arguments (not used) |
Value
The input pipeline object, returned invisibly. Called
for its side effect of printing detailed pipeline and model results.
Examples
pipe <- tl_pipeline(iris, Species ~ .)
summary(pipe)
Tidy Apriori Algorithm
Description
Mine association rules using the Apriori algorithm with tidy output
Usage
tidy_apriori(
transactions,
support = 0.01,
confidence = 0.5,
minlen = 2,
maxlen = 10,
target = "rules",
control = list(verbose = FALSE),
...
)
Arguments
transactions |
A transactions object or data frame |
support |
Minimum support (default: 0.01) |
confidence |
Minimum confidence (default: 0.5) |
minlen |
Minimum rule length (default: 2) |
maxlen |
Maximum rule length (default: 10) |
target |
Type of association mined: "rules" (default), "frequent itemsets", "maximally frequent itemsets" |
control |
A list of algorithmic controls for
|
... |
Further arguments passed to |
Value
A list of class "tidy_apriori" containing:
rules_tbl: tibble of rules, as
tidy_rulesreturns it, orNULLfor an itemset targetrules: original arules object, rules or itemsets
parameters: parameters used
n_rules: number of rules, or of itemsets for an itemset target
itemsets_tbl: for an itemset target only, a tibble with
itemset_id,itemset(its label),size, the quality measures, anditems, a list column holding each itemset's items
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
# Basic apriori
rules <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
# Access rules
rules$rules_tbl
}
Tidy CLARA (Clustering Large Applications)
Description
Performs CLARA clustering (scalable version of PAM)
Usage
tidy_clara(data, k, metric = "euclidean", samples = 50, sampsize = NULL, ...)
Arguments
data |
A data frame or tibble. CLARA samples observations, so it
takes no distance matrix; use |
k |
Number of clusters |
metric |
Distance metric (default: "euclidean") |
samples |
Number of samples to draw (default: 50) |
sampsize |
Sample size (default: min(n, 40 + 2*k)) |
... |
Further arguments passed to |
Value
A list of class "tidy_clara" containing:
clusters: tibble with observation IDs and cluster assignments
medoids: tibble of medoid values
silhouette_avg: average silhouette width
model: original
claraobject
Examples
# CLARA for large datasets
large_data <- iris[rep(1:nrow(iris), 10), 1:4]
clara_result <- tidy_clara(large_data, k = 3, samples = 50)
print(clara_result)
Cut Hierarchical Clustering Tree
Description
Cut dendrogram to obtain cluster assignments
Usage
tidy_cutree(hclust_obj, k = NULL, h = NULL)
Arguments
hclust_obj |
A tidy_hclust object or hclust object |
k |
Number of clusters (optional) |
h |
Height at which to cut (optional) |
Value
A tibble with columns .obs_id (observation identifier) and
cluster (integer cluster assignment).
Examples
hc <- tidy_hclust(USArrests, method = "ward.D2")
clusters <- tidy_cutree(hc, k = 3)
Tidy DBSCAN Clustering
Description
Performs density-based clustering with tidy output
Usage
tidy_dbscan(data, eps, minPts = 5, cols = NULL, distance = "euclidean")
Arguments
data |
A data frame, tibble, numeric matrix, or dist object |
eps |
Neighborhood radius (epsilon) |
minPts |
Minimum number of points to form a dense region (default: 5) |
cols |
Columns to include (tidy select).
If NULL, uses all numeric columns, or every column for
|
distance |
Distance metric if data is not a dist object (default:
"euclidean"): any method |
Value
A list of class "tidy_dbscan" containing:
clusters: tibble with observation IDs, cluster assignments (0 = noise), and the logical flags
is_noiseandis_coresummary: tibble with each cluster's size and number of core points
n_clusters: number of clusters (excluding noise)
n_noise: number of noise points
eps, minPts: the parameters used
model: original dbscan object
Examples
# Basic DBSCAN
db_result <- tidy_dbscan(iris, eps = 0.5, minPts = 5)
# With suggested eps from k-NN distance plot
eps_suggestion <- suggest_eps(iris, minPts = 5)
db_result <- tidy_dbscan(iris, eps = eps_suggestion$eps, minPts = 5)
Plot Dendrogram
Description
Create dendrogram visualization
Usage
tidy_dendrogram(hclust_obj, k = NULL, hang = 0.01, cex = 0.7)
Arguments
hclust_obj |
A tidy_hclust object or hclust object |
k |
Optional; number of clusters to highlight with rectangles |
hang |
Fraction of plot height to hang labels (default: 0.01) |
cex |
Label size (default: 0.7) |
Value
The hclust object, returned invisibly. The
dendrogram is plotted as a side effect.
Examples
hc <- tidy_hclust(USArrests, method = "ward.D2")
tidy_dendrogram(hc, k = 3)
Tidy Distance Matrix Computation
Description
Compute distance matrices with tidy output
Usage
tidy_dist(data, method = "euclidean", cols = NULL, ...)
Arguments
data |
A data frame or tibble |
method |
Character; distance method (default: "euclidean"). Options: "euclidean", "manhattan", "maximum", "gower" |
cols |
Columns to include (tidy select).
If NULL, uses all numeric columns, or every column for
|
... |
Additional arguments passed to distance functions |
Value
A dist object containing the computed
distance matrix.
Examples
d <- tidy_dist(iris[, 1:4], method = "euclidean")
Tidy Gap Statistic
Description
Compute gap statistic for determining optimal number of clusters
Usage
tidy_gap_stat(data, FUN_cluster = NULL, max_k = 10, B = 50, nstart = 25)
Arguments
data |
A data frame or tibble |
FUN_cluster |
Clustering function (default: uses kmeans internally) |
max_k |
Maximum number of clusters (default: 10) |
B |
Number of bootstrap samples (default: 50) |
nstart |
If using kmeans, number of random starts (default: 25) |
Value
A list of class "tidy_gap" containing:
gap_data: tibble with gap statistics for each k
k_firstSEmax: optimal k via
maxSE's firstSEmax method, the smallest k within one standard error of the first local maximum (most conservative)k_globalmax: optimal k via the globalmax method, the k with the largest gap (most liberal)
k_firstmax: optimal k via the firstmax method, the first local maximum of the gap
recommended_k: recommended k (uses firstSEmax)
model: the
clusGapresult
Examples
gap <- tidy_gap_stat(iris[, 1:4], max_k = 6, B = 10)
gap$recommended_k
Gower Distance Calculation
Description
Computes Gower distance for mixed data types (numeric, factor, ordered)
Usage
tidy_gower(data, weights = NULL)
Arguments
data |
A data frame or tibble |
weights |
Optional named vector of variable weights (default: equal weights) |
Details
Gower distance handles mixed data types:
Numeric: range-normalized Manhattan distance
Factor/Character: 0 if same, 1 if different
Ordered: treated as numeric ranks
Formula: d_ij = sum(w_k * d_ijk) / sum(w_k) where d_ijk is the dissimilarity for variable k between obs i and j
Value
A dist object containing Gower distances, with
the method attribute set to "gower". A pair of rows with
no variable observed in both has no defined distance and is NA,
as in daisy.
Examples
# Create example data with mixed types
car_data <- data.frame(
horsepower = c(130, 250, 180),
weight = c(1200, 1650, 1420),
color = factor(c("red", "black", "blue"))
)
# Compute Gower distance
gower_dist <- tidy_gower(car_data)
Tidy Hierarchical Clustering
Description
Performs hierarchical clustering with tidy output
Usage
tidy_hclust(data, method = "average", distance = "euclidean", cols = NULL)
Arguments
data |
A data frame, tibble, or dist object |
method |
Agglomeration method: "ward.D2", "single", "complete", "average" (default), "mcquitty", "median", "centroid" |
distance |
Distance metric if data is not a dist object (default: "euclidean") |
cols |
Columns to include (tidy select).
If NULL, uses all numeric columns, or every column for
|
Value
A list of class "tidy_hclust" containing:
model: hclust object
dist: distance matrix used
method: linkage method used
distance_method: distance metric used
data: original data (for plotting)
Examples
# Basic hierarchical clustering
hc_result <- tidy_hclust(USArrests, method = "average")
# With specific distance
hc_result <- tidy_hclust(mtcars, method = "complete", distance = "manhattan")
Tidy K-Means Clustering
Description
Performs k-means clustering with tidy output
Usage
tidy_kmeans(
data,
k,
cols = NULL,
nstart = 25,
iter_max = 100,
algorithm = "Hartigan-Wong"
)
Arguments
data |
A data frame or tibble |
k |
Number of clusters |
cols |
Columns to include (tidy select). If NULL, uses all numeric columns. |
nstart |
Number of random starts (default: 25) |
iter_max |
Maximum iterations (default: 100) |
algorithm |
K-means algorithm: "Hartigan-Wong" (default), "Lloyd", "Forgy", "MacQueen" |
Value
A list of class "tidy_kmeans" containing:
clusters: tibble with observation IDs and cluster assignments
centers: tibble of cluster centers
metrics: tibble with clustering quality metrics
sizes: integer vector of cluster sizes
model: original kmeans object
Examples
# Basic k-means
km_result <- tidy_kmeans(iris, k = 3)
Compute k-NN Distances
Description
Calculate distances to k-th nearest neighbor for each point
Usage
tidy_knn_dist(data, k = 4, cols = NULL)
Arguments
data |
A data frame or matrix |
k |
Number of nearest neighbors (default: 4) |
cols |
Columns to include (tidy select). If NULL, uses all numeric columns. |
Value
A tibble with columns .obs_id (observation identifier),
knn_dist (distance to k-th nearest neighbor), and rank
(rank of the k-NN distance).
Examples
knn <- tidy_knn_dist(iris[, 1:4], k = 5)
Tidy Multidimensional Scaling
Description
Unified interface for MDS methods with tidy output
Usage
tidy_mds(data, method = "classical", ndim = 2, distance = "euclidean", ...)
Arguments
data |
A data frame, tibble, or distance matrix |
method |
Character; "classical" (default), "metric", "nonmetric", "sammon", or "kruskal" |
ndim |
Number of dimensions for output (default: 2) |
distance |
Character; distance metric if data is
not already a dist object (default: "euclidean"): any method
|
... |
Additional arguments passed to specific MDS functions |
Value
A list of class "tidy_mds" containing:
config: tibble of MDS configuration (coordinates)
stress: goodness-of-fit measure (if applicable)
method: character string of method used
model: original model object
Examples
# Classical MDS
mds_result <- tidy_mds(eurodist, method = "classical")
print(mds_result)
Classical (Metric) MDS
Description
Performs classical multidimensional scaling using cmdscale()
Usage
tidy_mds_classical(dist_mat, ndim = 2, add_rownames = TRUE)
Arguments
dist_mat |
A distance matrix (dist object) |
ndim |
Number of dimensions (default: 2) |
add_rownames |
Preserve row names from distance matrix (default: TRUE) |
Value
A list of class "tidy_mds" containing:
config: tibble of MDS coordinates
stress:
NA(not applicable for classical MDS)gof: goodness-of-fit (proportion of variance retained)
eigenvalues: numeric vector of eigenvalues
method:
"Classical MDS"model: the
cmdscaleresult
Examples
d <- dist(USArrests)
mds <- tidy_mds_classical(d)
print(mds)
Kruskal's Non-metric MDS
Description
Performs Kruskal's isoMDS
Usage
tidy_mds_kruskal(dist_mat, ndim = 2, ...)
Arguments
dist_mat |
A distance matrix (dist object) |
ndim |
Number of dimensions (default: 2) |
... |
Additional arguments passed to MASS::isoMDS() |
Value
A list of class "tidy_mds" containing:
config: tibble of MDS coordinates
stress: Kruskal stress value
method:
"Kruskal's isoMDS"model: the
isoMDSresult
Examples
d <- dist(USArrests)
mds <- tidy_mds_kruskal(d)
Sammon Mapping
Description
Performs Sammon's non-linear mapping
Usage
tidy_mds_sammon(dist_mat, ndim = 2, ...)
Arguments
dist_mat |
A distance matrix (dist object) |
ndim |
Number of dimensions (default: 2) |
... |
Additional arguments passed to MASS::sammon() |
Value
A list of class "tidy_mds" containing:
config: tibble of MDS coordinates
stress: Sammon stress value
method:
"Sammon Mapping"model: the
sammonresult
Examples
d <- dist(USArrests)
mds <- tidy_mds_sammon(d)
SMACOF MDS (Metric or Non-metric)
Description
Performs MDS using SMACOF algorithm from the smacof package
Usage
tidy_mds_smacof(dist_mat, ndim = 2, type = "ratio", ...)
Arguments
dist_mat |
A distance matrix (dist object) |
ndim |
Number of dimensions (default: 2) |
type |
Character; "ratio" for metric, "ordinal" for non-metric (default: "ratio") |
... |
Additional arguments passed to smacof::mds() |
Value
A list of class "tidy_mds" containing:
config: tibble of MDS coordinates
stress: stress value from the SMACOF algorithm
method: character string describing the MDS type
model: the
mdsresult
Examples
d <- dist(USArrests)
mds <- tidy_mds_smacof(d, type = "ratio")
Tidy PAM (Partitioning Around Medoids)
Description
Performs PAM clustering with tidy output
Usage
tidy_pam(data, k, metric = "euclidean", cols = NULL, ...)
Arguments
data |
A data frame, tibble, or dist object |
k |
Number of clusters |
metric |
Distance metric (default: "euclidean"). Use "gower" for mixed data types. |
cols |
Columns to include (tidy select). If NULL, uses all columns. |
... |
Further arguments passed to |
Value
A list of class "tidy_pam" containing:
clusters: tibble with observation IDs and cluster assignments
medoids: tibble with one row per cluster: the medoid's row position in the data (
medoid_index, an integer) and, unlessdatawas a dist object, its valuessilhouette_avg: average silhouette width
silhouette_data: the silhouette information
pam()returns (itssilinfo)model: original pam object
Examples
# PAM with Euclidean distance
pam_result <- tidy_pam(iris, k = 3)
# PAM with Gower distance for mixed data
pam_result <- tidy_pam(mtcars, k = 3, metric = "gower")
Tidy Principal Component Analysis
Description
Performs PCA on a dataset using tidyverse principles. Returns a tidy list containing scores, loadings, variance explained, and the original model.
Usage
tidy_pca(data, cols = NULL, scale = TRUE, center = TRUE, method = "prcomp")
Arguments
data |
A data frame or tibble |
cols |
Columns to include in PCA (tidy select syntax). If NULL, uses all numeric columns. |
scale |
Logical; should variables be scaled to unit variance? Default TRUE. |
center |
Logical; should variables be centered? Default TRUE.
|
method |
Character; "prcomp" (default, recommended) or "princomp" |
Value
A list of class "tidy_pca" containing:
scores: tibble of PC scores with observation identifiers
loadings: tibble of variable loadings in long format
variance: tibble of variance explained by each PC
model: the original prcomp/princomp object
settings: list of scale, center, method used
Examples
# Basic PCA
pca_result <- tidy_pca(USArrests)
# Access components
pca_result$scores
pca_result$loadings
pca_result$variance
Create PCA Biplot
Description
Visualize both observations and variables in PC space
Usage
tidy_pca_biplot(
pca_obj,
pc_x = 1,
pc_y = 2,
color_by = NULL,
arrow_scale = 1,
label_obs = FALSE,
label_vars = TRUE
)
Arguments
pca_obj |
A tidy_pca object |
pc_x |
Principal component for x-axis (default: 1) |
pc_y |
Principal component for y-axis (default: 2) |
color_by |
Optional grouping to colour points by: either a column name present in the PCA scores, or a vector as long as the data. |
arrow_scale |
Scaling factor for variable arrows (default: 1) |
label_obs |
Logical; label observations? (default: FALSE) |
label_vars |
Logical; label variables? (default: TRUE) |
Value
A ggplot object.
Examples
pca <- tidy_pca(USArrests)
tidy_pca_biplot(pca)
Create PCA Scree Plot
Description
Visualize variance explained by each principal component
Usage
tidy_pca_screeplot(pca_obj, type = "proportion", add_line = TRUE)
Arguments
pca_obj |
A tidy_pca object |
type |
Character; "variance" or "proportion" (default) |
add_line |
Logical; add horizontal line at eigenvalue = 1? (for Kaiser criterion) |
Value
A ggplot object.
Examples
pca <- tidy_pca(USArrests)
tidy_pca_screeplot(pca)
Convert Association Rules to Tidy Tibble
Description
Convert Association Rules to Tidy Tibble
Usage
tidy_rules(rules)
Arguments
rules |
A rules object from arules |
Value
A tibble with columns rule_id, lhs, rhs,
the quality measures (e.g., support, confidence,
lift), and the list columns lhs_items and
rhs_items, each rule's items on that side as a character
vector. The lhs and rhs labels are for reading; the
helpers that match items read the lists, since an item name can hold
the comma that separates items in a label. An empty rule set gives a
zero-row tibble with the same columns.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
rules_obj <- arules::apriori(Groceries,
parameter = list(supp = 0.001, conf = 0.5))
rules_tbl <- tidy_rules(rules_obj)
}
Tidy Silhouette Analysis
Description
Compute silhouette statistics for cluster validation
Usage
tidy_silhouette(clusters, dist_mat)
Arguments
clusters |
Vector of cluster assignments: numeric, factor or
character. Numeric labels, or labels that read as whole numbers (such
as the |
dist_mat |
Distance matrix (dist object) |
Value
A list of class "tidy_silhouette" containing:
silhouette_data: tibble with silhouette values for each observation
avg_width: average silhouette width
cluster_avg: average silhouette width by cluster
Examples
km <- kmeans(iris[, 1:4], centers = 3, nstart = 25)
d <- dist(iris[, 1:4])
sil <- tidy_silhouette(km$cluster, d)
Silhouette Analysis Across Multiple k Values
Description
Silhouette Analysis Across Multiple k Values
Usage
tidy_silhouette_analysis(
data,
max_k = 10,
method = "kmeans",
nstart = 25,
dist_method = "euclidean",
linkage_method = "average"
)
Arguments
data |
A data frame or tibble |
max_k |
Maximum number of clusters to test (default: 10) |
method |
Clustering method: "kmeans" (default) or "hclust" |
nstart |
If kmeans, number of random starts (default: 25) |
dist_method |
Distance metric (default: "euclidean") |
linkage_method |
If hclust, linkage method (default: "average") |
Value
A tibble with columns k and avg_sil_width. The
"optimal_k" attribute contains the k with the highest average
silhouette width.
Examples
sil_analysis <- tidy_silhouette_analysis(iris[, 1:4], max_k = 6)
Classification Functions for tidylearn
Description
Logistic regression and classification metrics functionality
Cloud data-egress consent for tidylearn
Description
The consent gate that stands between a compute = "cloud" call and any training data leaving the user's machine.
Cost controls for tidylearn cloud compute
Description
Bounding what a cloud fit can cost, and keeping in-flight jobs visible.
Cloud endpoint resolution for tidylearn
Description
Resolving and validating the Modal Web Function endpoint
that compute = "cloud" submits to.
Model serialisation for tidylearn cloud compute
Description
Converting a fitted tidylearn model to bytes and back, so it can cross the boundary between a remote worker and the caller's R session.
Coefficient Inference for tidylearn
Description
Coefficients, standard errors and confidence intervals as a
tibble, for the methods that have them. tl_table_coefficients()
formats the same numbers as a gt table.
tidylearn: A Unified Tidy Interface to R's Machine Learning Ecosystem
Description
Core functionality for tidylearn. This package provides a unified tidyverse-compatible interface to established R machine learning packages including glmnet, randomForest, xgboost, e1071, rpart, gbm, nnet, cluster, and dbscan. The underlying algorithms are unchanged - tidylearn wraps them with consistent function signatures, tidy tibble output, and unified ggplot2-based visualization. Supervised models keep the wrapped object at model$fit; unsupervised ones put it at model$fit$model, alongside the tidied components.
Deep Learning for tidylearn
Description
Deep learning functionality using Keras/TensorFlow
Advanced Diagnostics Functions for tidylearn
Description
Functions for advanced model diagnostics, assumption checking, and outlier detection
Interaction Analysis Functions for tidylearn
Description
Functions for testing, visualizing, and analyzing interactions
Metrics Functionality for tidylearn
Description
Functions for calculating model evaluation metrics
Model Selection Functions for tidylearn
Description
Functions for stepwise model selection, cross-validation, and hyperparameter tuning
Neural Networks for tidylearn
Description
Neural network functionality for classification and regression
Model Pipeline Functions for tidylearn
Description
Functions for creating end-to-end model pipelines
Data Reading Functions for tidylearn
Description
Functions for reading data from diverse sources into tidy
tidylearn_data objects. The main dispatcher
tl_read() auto-detects the format from the file
extension and routes to the appropriate reader.
All readers return a tidylearn_data object,
which is a tibble subclass carrying metadata about
the data source.
Details
Supported file formats:
-
CSV:
.csvfiles via readr (with base R fallback), and.txtfiles named directly -
TSV:
.tsvfiles via readr (with base R fallback) -
Excel:
.xls,.xlsx,.xlsmfiles via readxl -
Parquet:
.parquetfiles via nanoparquet -
JSON:
.jsonfiles, and newline-delimited.ndjsonfiles, via jsonlite -
RDS:
.rdsfiles via basereadRDS() -
RData:
.rdata,.rdafiles via baseload()
CSV and TSV files compressed with gzip, bzip2 or xz
(data.csv.gz) are recognised by the extension under the
compression one.
Supported databases (via DBI):
-
SQLite:
.sqlite,.dbfiles via RSQLite -
PostgreSQL: via RPostgres
-
MySQL/MariaDB: via RMariaDB
-
BigQuery:
bigquery://project/datasetURIs via bigrquery
Supported cloud/API sources:
-
S3:
s3://URIs via paws.storage -
GitHub: raw file download from repositories
-
Kaggle: dataset download via Kaggle CLI
A file:// URL is read as the local path it names. Other web
URLs, and URLs with any other scheme such as ftp://, are not
read; download the file first.
Multi-file reading:
-
Multiple paths: pass a character vector to
tl_read() -
Directories:
tl_read_dir()scans for data files with optional pattern/format filtering and recursive scanning -
Zip archives:
tl_read_zip()extracts and reads from.zipfiles
When combining multiple files, a source_file column is added to
identify the origin of each row: the file's path below the directory
or archive it came from, or, for paths given directly, below the
deepest folder they share. Files in one folder are labelled by their
bare names.
Directory and archive scans read the extensions listed above except
.txt, which in a folder is as likely to hold notes as data.
Name a .txt file directly, or select it with pattern,
to read it.
Data Reading Backends for tidylearn
Description
Backend readers for databases and cloud/API sources.
All backends are optional dependencies checked at call time via
tl_check_packages().
Details
Database backends (via DBI):
-
SQLite: via RSQLite
-
PostgreSQL: via RPostgres
-
MySQL/MariaDB: via RMariaDB
-
BigQuery: via bigrquery
Cloud/API backends:
-
S3: via paws.storage
-
GitHub: via base
download.file() -
Kaggle: via Kaggle CLI
Regression Functions for tidylearn
Description
Linear and polynomial regression functionality
Regularization Functions for tidylearn
Description
Ridge, Lasso, and Elastic Net regularization functionality
Support Vector Machines for tidylearn
Description
SVM functionality for classification and regression
Table Functions for tidylearn
Description
Functions for producing formatted gt tables
from tidylearn models. Provides a parallel interface to
the plot functions: tl_table(model, type)
dispatches to the appropriate table formatter based on model type.
Requires the gt package (suggested dependency).
Tree-based Methods for tidylearn
Description
Decision trees, random forests, and boosting functionality
Hyperparameter Tuning Functions for tidylearn
Description
Functions for automatic hyperparameter tuning and selection
Visualization Functions for tidylearn
Description
General visualization functions for tidylearn models
High-Level Workflows for Common Machine Learning Patterns
Description
Functions providing end-to-end workflows that showcase tidylearn's ability to seamlessly combine multiple learning paradigms
XGBoost Functions for tidylearn
Description
XGBoost-specific implementation for gradient boosting
Cluster-Based Features
Description
Add cluster assignments as features for supervised learning. This semi-supervised approach can capture non-linear patterns.
Usage
tl_add_cluster_features(data, response = NULL, method = "kmeans", ...)
Arguments
data |
A data frame |
response |
Response variable name (will be excluded from clustering) |
method |
Clustering method: "kmeans", "pam", "hclust", "dbscan" |
... |
Additional arguments for clustering |
Value
The original data frame with an additional factor column named
cluster_<method> containing cluster assignments. The fitted
cluster model is stored as an attribute "cluster_model".
Examples
# Add cluster features before supervised learning
data_with_clusters <- tl_add_cluster_features(iris, response = "Species",
method = "kmeans", k = 3)
model <- tl_model(data_with_clusters, Species ~ ., method = "forest")
Anomaly-Aware Supervised Learning
Description
Detect outliers using DBSCAN or other methods, then optionally remove them or down-weight them before supervised learning.
Usage
tl_anomaly_aware(
data,
formula,
response,
anomaly_method = "dbscan",
action = "flag",
supervised_method = "tree",
...
)
Arguments
data |
A data frame |
formula |
Model formula |
response |
Response variable name, left out of the detection |
anomaly_method |
Method for anomaly detection. Only "dbscan" is implemented; its noise points are the anomalies. |
action |
Action to take: "remove", "flag", "downweight".
|
supervised_method |
Supervised learning method (default:
|
... |
Additional arguments for DBSCAN, such as |
Details
DBSCAN runs on the predictors the formula names, on their own scale: its
eps and minPts (defaults 0.5 and 5, passed through
...) are a distance and a count in those units. If every row comes
out as noise the call stops, since no normal data would be left to model.
Value
A tidylearn model object with additional class
"tidylearn_anomaly_aware". The model includes an
anomaly_info element with anomaly_model,
is_anomaly (logical vector), n_anomalies, and
action.
Examples
model <- tl_anomaly_aware(iris, Species ~ ., response = "Species",
anomaly_method = "dbscan", action = "flag")
Find important interactions automatically
Description
Find important interactions automatically
Usage
tl_auto_interactions(
data,
formula,
top_n = 3,
min_r2_change = 0.01,
max_p_value = 0.05,
exclude_vars = NULL
)
Arguments
data |
A data frame containing the data |
formula |
A formula specifying the base model without interactions |
top_n |
Number of top interactions to return, a whole number. With
|
min_r2_change |
Minimum change in R-squared to consider |
max_p_value |
Maximum p-value for significance |
exclude_vars |
Character vector of predictor variables that may not
appear in a selected interaction. They stay in the model as main effects.
Every name must be a predictor in |
Value
A tidylearn model object (class "tidylearn_model") fitted
with the top significant interaction terms added to the formula.
The interaction test results and selected interactions are stored as
attributes "interaction_tests" and
"selected_interactions", data frames in the layout
tl_test_interactions returns. Both are present when no
interaction is added, with no rows where there is nothing to report.
Examples
model <- tl_auto_interactions(mtcars, mpg ~ wt + hp + cyl, top_n = 2)
Auto ML: Automated Machine Learning Workflow
Description
Automatically explores multiple modeling approaches including dimensionality reduction, clustering, and various supervised methods. Returns the best performing model, scored by cross-validation where the time budget allows.
Usage
tl_auto_ml(
data,
formula,
task = "auto",
use_reduction = TRUE,
use_clustering = TRUE,
time_budget = 300,
cv_folds = 5,
metric = NULL
)
Arguments
data |
A data frame |
formula |
Model formula (for supervised learning) |
task |
Task type: "classification", "regression", or "auto"
(default), which takes it from the response the formula computes, so
|
use_reduction |
Whether to try dimensionality reduction (default: TRUE) |
use_clustering |
Whether to add cluster features (default: TRUE) |
time_budget |
Time budget in seconds (default: 300). The budget is checked between model fits, not during them: once a model starts training it runs to completion, because R cannot safely interrupt C-level code (randomForest, xgboost, e1071). A run can therefore overshoot the budget by the length of the last fit it started. The budget gates the workflow as follows:
The example below, with |
cv_folds |
Number of cross-validation folds (default: 5). Reducing this (e.g. to 2 or 3) is an effective way to stay closer to the time budget since CV is typically the most expensive step. |
metric |
Evaluation metric (default: "accuracy" for classification, "rmse" for regression). Classification takes "accuracy", "precision", "recall", "sensitivity", "specificity", "f1", "auc" or "pr_auc"; regression takes "rmse", "mse", "mae", "mape" or "rsq". It is checked before any model is fitted. |
Details
The PCA and cluster variants are built from the formula's predictors
only, so a column the formula leaves out (y ~ . - id) reaches no
candidate. The cluster variants add the cluster assignment to the
formula's terms.
Value
A list with class "tidylearn_automl" containing:
- best_model
The best tidylearn model object
- models
Named list of all successfully trained models
- leaderboard
Tibble ranking models by the chosen metric, with columns
model,scoreandevaluation. Theevaluationcolumn records how each score was obtained –"cv"for cross-validated,"train"for training-set metrics, which are optimistic. Scores of different kinds are not directly comparable; a mixed leaderboard means the budget ran short of cross-validating every model.- task
Detected or specified task type
- metric
Metric used for ranking
- runtime
Total elapsed time as a difftime object
Examples
# Quick run with fast models only (< 30s budget skips forest/SVM/XGBoost)
result <- tl_auto_ml(iris, Species ~ .,
time_budget = 10,
use_reduction = FALSE,
use_clustering = FALSE,
cv_folds = 2)
result$leaderboard
Calculate classification metrics
Description
Scores predicted classes, and optionally class probabilities, against observed classes.
Usage
tl_calc_classification_metrics(
actuals,
predicted,
predicted_probs = NULL,
metrics = c("accuracy", "precision", "recall", "f1", "auc"),
thresholds = NULL,
...
)
Arguments
actuals |
Observed classes: a factor, or a character, logical or numeric vector. |
predicted |
Predicted classes, one for each element of
|
predicted_probs |
Class probabilities, needed for |
metrics |
Character vector of metrics to compute, from
|
thresholds |
Optional numeric vector of cut-offs on the positive
class's probability, for binary classification. Each adds rows
scoring the classes that cut-off assigns. Needs
|
... |
Not used. |
Details
The classes, in order, are the levels of predicted when it is a
factor of two or more levels – which is how predict() returns
them, in the model's order – and otherwise the classes present in
actuals. Any other class found in actuals or
predicted follows them. For two classes the second is the
positive class. tl_evaluate instead leaves out rows of a
class the model was never trained on, since it knows the model's
classes.
A row missing its observed class, its prediction or one of its
probabilities is dropped before anything is computed, so every metric
describes the same rows. With no row left – actuals empty, or
every row incomplete – there is nothing to score, and it is an error
of class tidylearn_no_scored_rows.
Value
A tibble with columns metric (character)
and value (numeric), one row per requested metric. For more
than two classes, "auc" is followed by an auc_<class>
row for each class.
"auc" and "pr_auc" need at least two classes among the
scored rows, and are NA, with a warning, when there is only
one. A class with no scored row has no one-vs-rest area: its
auc_<class> row is NA, the averages cover the other
classes, and a warning names it.
With thresholds, six rows per cut-off follow – for a cut-off
of 0.5, accuracy_t0.5, precision_t0.5,
recall_t0.5, f1_t0.5, f2_t0.5 and
f0.5_t0.5 – and the tibble gains a threshold column,
NA on the other rows.
Examples
model <- tl_model(iris, Species ~ ., method = "forest")
preds <- predict(model)
tl_calc_classification_metrics(iris$Species, preds$.pred)
Calculate the area under the precision-recall curve
Description
Reads the area off a curve ROCR built, and agrees with
yardstick::pr_auc(). ROCR's curve starts at recall 0, where
nothing is yet called positive and precision is 0/0. The area has to
start there too, at precision 1 as yardstick's does: integrating from
the first finite point lost everything before it, so a perfect ranking
of 5 positives in 20 scored 0.8, a tree on two iris species 0.07, and
constant scores left no area at all.
Usage
tl_calculate_pr_auc(perf)
Arguments
perf |
A ROCR performance object of precision against recall |
Value
The area under the precision-recall curve, or NA when
the curve has fewer than two points
Check model assumptions
Description
Check model assumptions
Usage
tl_check_assumptions(model, test = TRUE, verbose = TRUE)
Arguments
model |
A tidylearn model object |
test |
Logical; whether to perform statistical tests |
verbose |
Logical; whether to print test results and explanations |
Value
A named list with one element per assumption checked
(linearity, independence, homoscedasticity,
normality, multicollinearity, outliers), each
containing assumption (character label), check (logical,
NA when the test could not decide, or NULL when no test
was run), details (character), and recommendation
(character). An additional overall element summarises the
number of assumptions checked, violated, and satisfied; an NA or
NULL check counts as neither.
Logistic regression assumes neither normal residuals nor a constant
variance, so for a logistic model normality and
homoscedasticity have a NULL check and a note saying so.
For a factor, multicollinearity is judged on GVIF^(1/Df), the
generalised VIF on the scale of an ordinary one.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_check_assumptions(model)
Detect local GPU availability for tidylearn methods
Description
Reports whether the local machine has a CUDA-capable GPU and which
tidylearn backends (xgboost, keras, tensorflow) are positioned
to use it. Detection is intentionally cheap: it parses nvidia-smi
output, giving up with a warning if nvidia-smi has not answered
within 10 seconds, and reads which R packages are installed from the
library without loading them, so it neither starts Python nor fits a
model. A backend reported as gpu_likely_works = TRUE may still fall
back to CPU if it was not compiled or configured with CUDA support —
confirm with a small real fit before relying on it for production
workloads.
Usage
tl_check_gpu(verbose = FALSE)
Arguments
verbose |
Logical. If |
Details
Apple MPS (Metal Performance Shaders) is intentionally not detected in this iteration; see the issue tracker for the MPS feature request.
Value
An object of class tidylearn_gpu_check: a list with components
any_gpu (logical), cuda (driver info list with driver_present,
device_count, device_names, driver_version), backends (per-backend
status list each containing installed, gpu_likely_works, notes),
and messages (character vector). A print() method is provided.
Examples
# Safe to call anywhere: probes for nvidia-smi on the PATH and checks
# which backend packages are installed. Reports no GPU rather than
# failing when there isn't one.
gpu <- tl_check_gpu()
gpu$any_gpu
if (gpu$any_gpu && gpu$backends$xgboost$gpu_likely_works) {
# xgboost with GPU is worth trying for this workload
}
Allow an additional host for cloud uploads in this R session
Description
tidylearn uploads training data only to Modal's own hosts
(*.modal.run, *.modal.com). Modal customers serving Web Functions
from a custom domain can add that domain here.
Usage
tl_cloud_allow_host(host)
Arguments
host |
A character vector of host names to allow, or |
Details
This widens the set of destinations your data may be sent to, so it
is deliberately a per-session call rather than an option or an
environment variable: a shared .Rprofile or an inherited environment
should not be able to add a destination without you writing the call.
Additions are never persisted and are forgotten when the session ends.
Hosts match themselves and their subdomains. Adding
"fits.example.com" accepts https://fits.example.com and
https://a.fits.example.com, and nothing else.
Give the endpoint's full host name. A name of fewer than three labels,
such as "example.com" or "co.uk", is refused: two labels can be a
public suffix, under which anyone can register a site, and allowing
one would admit all of them.
Value
The full allowlist after the change, invisibly.
See Also
tl_cloud_allowed_hosts(), and T9 in
system.file("security/threat-model.md", package = "tidylearn").
Examples
tl_cloud_allow_host("fits.example.com")
tl_cloud_allowed_hosts()
tl_cloud_allow_host(NULL)
Hosts tidylearn will currently upload to
Description
Modal's own hosts, plus anything added with tl_cloud_allow_host()
during this session.
Usage
tl_cloud_allowed_hosts()
Value
A character vector of host names.
See Also
Examples
tl_cloud_allowed_hosts()
Grant or revoke cloud upload consent for this R session
Description
Fitting with compute = "cloud" uploads your training data to your
own Modal account, which is a third party. tidylearn will not do that
without explicit consent on every call.
Usage
tl_cloud_consent(consent = TRUE)
Arguments
consent |
|
Details
There are two ways to give it. Pass confirm_upload = TRUE to each
tl_model() call, or call tl_cloud_consent() once to opt in for the
rest of the session. The session lock exists for batch and
non-interactive work, where a per-call argument is repetitive.
The lock is not persisted. It is forgotten when the R session
ends, and it is never written to disk. Revoke it early with
tl_cloud_consent(FALSE).
tidylearn never prompts interactively for consent, so scripts, CI and
Rscript behave the same as an interactive session.
Value
The previous consent state, invisibly.
See Also
The full contract is in
system.file("security/threat-model.md", package = "tidylearn").
Examples
# Opt in for the session, then revoke.
old <- tl_cloud_consent(TRUE)
tl_cloud_consent(FALSE)
Cloud jobs submitted in this R session
Description
Every cloud submission is recorded here so that no job runs invisibly. A job disappears from this list when its result is collected or it is cancelled.
Usage
tl_cloud_jobs()
Details
A submitted job runs on Modal whatever this R session does. If the session ends while a job is still listed here, that job keeps running and keeps billing until it finishes or its timeout kills it — this list cannot survive the session, but the timeout does.
Value
A tibble with one row per in-flight job: call_id, method,
submitted_at, timeout_seconds and worst_case_cost. Zero rows
when nothing is in flight.
Examples
# Nothing in flight in a fresh session
tl_cloud_jobs()
Model coefficients as a tibble
Description
Returns the coefficients of a fitted tidylearn model as a tibble, with
standard errors, test statistics, p-values and – on request – confidence
intervals. Available for "linear", "polynomial",
"logistic", "ridge", "lasso" and
"elastic_net". Other methods have no coefficients; use
tl_table_importance for those.
Usage
tl_coefficients(
model,
conf_int = FALSE,
level = 0.95,
exponentiate = FALSE,
lambda = "1se"
)
Arguments
model |
A tidylearn supervised model object from
|
conf_int |
Whether to add |
level |
Confidence level for the interval (default 0.95), a number
strictly between 0 and 1. Used only when |
exponentiate |
Whether to report |
lambda |
For regularised methods: |
Details
Intervals are Wald intervals, computed from the same standard errors as
the statistic and p_value columns beside them, so the
interval and the p-value always agree about whether zero is excluded. For
"linear" and "polynomial" that means t quantiles on
the residual degrees of freedom, which is exactly what
confint returns for an lm. For
"logistic" it means z quantiles, which is what the reported
z statistic implies but not what confint() gives – that
profiles the likelihood, which is the better interval when the sample is
small or a class is nearly separated. Call
stats::confint(model$fit) when you want it.
A rank-deficient fit – two perfectly collinear predictors, or an
interaction of factors with a combination no row has – cannot estimate
every term. Those terms are
returned with an NA estimate rather than dropped, so a term named
in the formula never disappears from the output without saying so.
Value
A tibble, one row per model term. For "linear",
"polynomial" and "logistic": term,
estimate, std_error, statistic, p_value,
plus conf_low and conf_high when conf_int = TRUE.
For regularised methods: term, estimate and the
lambda the estimate came from – glmnet reports no standard
errors, so there is nothing to test or bound. A multiclass
regularised model has one set of coefficients per class, so its rows
are led by a class column; exponentiate is not available
for it, because glmnet's multinomial coefficients are not relative to
a reference class.
See Also
tl_table_coefficients for the same numbers as a
formatted table.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_coefficients(model)
tl_coefficients(model, conf_int = TRUE)
# Odds ratios from a logistic fit
am_data <- transform(mtcars, am = factor(am))
model <- tl_model(am_data, am ~ wt, method = "logistic")
tl_coefficients(model, conf_int = TRUE, exponentiate = TRUE)
Compare models using cross-validation
Description
Each model is refitted on every fold from its formula, method and
fitting arguments. A model a refit would not reproduce is refused: one
built by tl_semisupervised or tl_anomaly_aware,
or one of tl_auto_ml's candidates fitted on features it
engineered.
Usage
tl_compare_cv(data, models, folds = 5, metrics = NULL, ...)
Arguments
data |
A data frame containing the training data |
models |
A named list of supervised tidylearn models, all of them
classification or all regression. An unnamed model is named
|
folds |
Number of cross-validation folds, a whole number between 2
and |
metrics |
Character vector of metrics to compute, from those
|
... |
Arguments passed to |
Value
A list with two elements:
$fold_metricsA data frame with columns
metric,value,fold, andmodelcontaining per-fold results for every model. A metric undefined on a fold –"auc"on a fold holding one class – isNAthere, and so is every metric of a fold none of whose rows can be scored, with a warning naming the fold.$summaryA data frame with columns
model,metric,mean_value,sd_value,min_value, andmax_valuesummarizing cross-validation performance over the folds with a value. A metric with no value on any fold isNAthroughout.
Examples
m1 <- tl_model(mtcars, mpg ~ wt, method = "linear")
m2 <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
cv <- tl_compare_cv(mtcars, list(simple = m1, full = m2), folds = 3)
cv$summary
Compare models from a pipeline
Description
Compare models from a pipeline
Usage
tl_compare_pipeline_models(pipeline, metrics = NULL)
Arguments
pipeline |
A tidylearn pipeline object with results |
metrics |
Character vector of metrics to compare, each one the pipeline scored (if NULL, uses all available) |
Value
A ggplot object showing a faceted bar
chart comparing metric values across models, with the best model
highlighted.
Examples
pipe <- tl_pipeline(iris, Species ~ .,
models = list(
tree = list(method = "tree"),
forest = list(method = "forest", ntree = 100)
),
evaluation = list(validation = "cv", cv_folds = 3))
pipe <- tl_run_pipeline(pipe, verbose = FALSE)
tl_compare_pipeline_models(pipe)
# Restrict the comparison to one metric
tl_compare_pipeline_models(pipe, metrics = "accuracy")
Advise on the best compute tier for a tidylearn fit
Description
Estimates runtime and feasibility on local CPU, local GPU (when available and applicable), and cloud GPU (stubbed until cloud integration lands), then returns a structured recommendation. Useful before kicking off a long fit — call this first to see whether the problem is laptop-sized, GPU-sized, or cloud-sized.
Usage
tl_compute_advisor(x, ...)
## S3 method for class 'character'
tl_compute_advisor(
x,
data,
formula = NULL,
hyperparams = list(),
gpu_check = NULL,
...
)
## S3 method for class 'tidylearn_supervised'
tl_compute_advisor(
x,
data = NULL,
formula = NULL,
hyperparams = list(),
gpu_check = NULL,
...
)
## Default S3 method:
tl_compute_advisor(x, ...)
Arguments
x |
Either a method name (character scalar — same names as
accepted by |
... |
Unused, reserved for method-specific extensions. |
data |
A data frame. Required when |
formula |
Optional formula. Used to determine the number of
effective predictors: the terms it expands to against |
hyperparams |
Named list of hyperparameters that affect runtime:
|
gpu_check |
Optional |
Details
Estimates are deliberately rough — order-of-magnitude, not bills. Per-method scaling constants are calibrated against typical hardware and will be off by 2-3x in either direction for any individual job. Treat the recommendation as a starting point, not gospel.
Value
An object of class tidylearn_compute_advice containing
problem, local_cpu, local_gpu, cloud, recommendation,
and reasoning. A print() method is provided.
Examples
# Estimating from a method name needs neither a GPU nor the backend
# package -- it is arithmetic over the problem dimensions
advice <- tl_compute_advisor("xgboost", iris, Species ~ .,
hyperparams = list(nrounds = 1000))
advice$recommendation
print(advice)
# Dispatching on a fitted model requires the backend to be installed.
# CRAN asks examples to use at most two threads.
if (requireNamespace("xgboost", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "xgboost", nthread = 2)
tl_compute_advisor(model)
}
Cross-validation for tidylearn models
Description
Each fold's model is scored with tl_evaluate, so the
response is read as that function reads it: a transformed left-hand
side on its own scale, and classes against the fold model's.
Usage
tl_cv(data, formula, method, folds = 5, metrics = NULL, transform = NULL, ...)
Arguments
data |
Data frame |
formula |
Model formula |
method |
Modeling method |
folds |
Number of cross-validation folds, a whole number between 2
and |
metrics |
Character vector of metrics to compute on each fold,
passed to |
transform |
Optional function for feature engineering that has to
be refitted per fold. It is called with the training rows of each
fold and must return a list with an |
... |
Additional arguments passed to |
Value
A list with two elements:
$foldsA list of per-fold evaluation tibbles, each with
metricandvaluecolumns.$summaryA tibble with columns
metric,mean, andsdsummarizing performance across folds. A metric undefined on a fold – auc on a fold holding one class – isNAthere and left out of the mean and sd. So is every metric of a fold none of whose rows can be scored, with a warning giving the reason. A metric with no value on any fold hasNAmean and sd.
Examples
cv <- tl_cv(mtcars, mpg ~ wt + hp, method = "linear", folds = 3)
cv$summary
Create interactive visualization dashboard for a model
Description
Create interactive visualization dashboard for a model
Usage
tl_dashboard(model, new_data = NULL, ...)
Arguments
model |
A tidylearn model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
... |
Additional arguments |
Value
A shinyApp object.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
app <- tl_dashboard(model)
Create pre-defined parameter grids for common models
Description
Create pre-defined parameter grids for common models
Usage
tl_default_param_grid(method, size = "medium", is_classification = TRUE)
Arguments
method |
Model method ("tree", "forest", "boost", "svm", "xgboost", etc.) |
size |
Grid size: "small", "medium", "large" |
is_classification |
Whether the grid is for a classification task
(the default) or a regression one. For regression, an |
Value
A named list of parameter values suitable for passing to
tl_tune_grid or tl_tune_random. Each
element is a numeric or character vector of candidate values for
that hyperparameter, or for "deep"'s hidden_layers a list
of layer-size vectors. The grid is built without the data, so a
"forest" mtry can exceed the number of predictors; the
tuners cap it. "polynomial" tunes degree. The
"xgboost" grids draw on the values tl_tune_xgboost
searches by default, and add nrounds, which that function
chooses by early stopping and tl_tune_grid has to tune.
"linear" and "logistic" have no tuneable
hyperparameter and return an empty list with a warning, as does an
unknown method.
Examples
grid <- tl_default_param_grid("tree", size = "small")
grid <- tl_default_param_grid("forest", size = "medium")
grid <- tl_default_param_grid("svm", is_classification = FALSE)
Detect outliers in the data
Description
Detect outliers in the data
Usage
tl_detect_outliers(
data,
variables = NULL,
method = "iqr",
threshold = NULL,
plot = TRUE
)
Arguments
data |
A data frame containing the data |
variables |
Character vector of variables to check for outliers |
method |
Method for outlier detection: "boxplot", "z-score", "cook", "iqr", "mahalanobis" |
threshold |
Threshold for outlier detection |
plot |
Logical; whether to create a plot of outliers |
Value
A list with outlier detection results:
- method
The detection method used (character).
- method_name
Human-readable method name (character).
- threshold
The threshold value used (numeric).
- threshold_label
Formatted threshold description (character).
- outlier_flags
A logical matrix (observations x variables),
NAwhere a value is missing. For"cook"and"mahalanobis"a row's flags are the same in every column, andNAwhen any of its values is missing.- any_outlier
Logical vector indicating if each observation is an outlier in any variable. Missing flags are ignored, so a row with no flag at all is
FALSE.- outlier_counts
List with
total,by_variable, andby_observationcounts.- outlier_indices
Integer vector of outlier row indices.
- plot
A
ggplotobject, orNULLifplot = FALSE.
Examples
tl_detect_outliers(mtcars, variables = c("mpg", "wt"), method = "iqr")
Create a comprehensive diagnostic dashboard
Description
Create a comprehensive diagnostic dashboard
Usage
tl_diagnostic_dashboard(
model,
include_influence = TRUE,
include_assumptions = TRUE,
include_performance = TRUE,
arrange_plots = "grid"
)
Arguments
model |
A tidylearn model object whose fit is an |
include_influence |
Logical; whether to include influence diagnostics |
include_assumptions |
Logical; whether to include assumption checks |
include_performance |
Logical; whether to include performance metrics |
arrange_plots |
Layout arrangement (e.g., "grid", "row", "column") |
Value
A grid.arrange object (a
grob) containing the arranged diagnostic plots.
Examples
if (requireNamespace("gridExtra")) {
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_diagnostic_dashboard(model)
}
Evaluate a tidylearn model
Description
Scores a supervised model's predictions against the observed response.
Usage
tl_evaluate(object, new_data = NULL, metrics = NULL, ...)
Arguments
object |
A tidylearn model object |
new_data |
Optional new data for evaluation (if NULL, uses training data) |
metrics |
Character vector of metrics to compute. If |
... |
Additional arguments passed to |
Details
The observed response is the formula's left-hand side evaluated on the
scored rows, so a model of log(mpg) is scored against
log(mpg), the scale it predicts on. For classification, the
observed classes are read against the classes the model was trained on,
whose second is the positive class; rows of a class the model never saw
are left out with a warning. Rows missing the response or a prediction
are dropped. tl_calc_classification_metrics describes
"auc" and "pr_auc" when the scored rows hold a single
class or lack one.
With no row left to score – new_data has no rows, or every row
is dropped – it is an error of class tidylearn_no_scored_rows,
naming the reason. tl_cv catches that class and leaves the
fold out.
Value
A tibble with columns metric (character)
and value (numeric), containing one row per requested metric,
and for "auc" on more than two classes an auc_<class>
row per class as well. An unsupervised model has no response to score
against and returns the single row metric = "completed",
value = 1.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_evaluate(model)
tl_evaluate(model, metrics = c("rmse", "mape"))
Evaluate metrics at different thresholds
Description
Evaluate metrics at different thresholds
Usage
tl_evaluate_thresholds(actuals, probs, thresholds, pos_class)
Arguments
actuals |
Actual values (ground truth) |
probs |
Predicted probabilities |
thresholds |
Vector of thresholds to evaluate |
pos_class |
The positive class |
Value
A tibble of metrics at different thresholds
Positive-class argument for yardstick binary metrics
Description
tidylearn's positive class is the second factor level. yardstick's
default is the first, and its event_level argument only applies
to the binary case – passing it for a multiclass problem warns.
Usage
tl_event_level_args(actuals)
Arguments
actuals |
A factor of ground-truth values |
Value
A list to splice into a yardstick call: event_level =
"second" for a two-level factor, empty otherwise
Exploratory Data Analysis Workflow
Description
Comprehensive EDA combining unsupervised learning techniques to understand data structure before modeling
Usage
tl_explore(data, response = NULL, max_components = 5, k_range = 2:6)
Arguments
data |
A data frame |
response |
Optional response variable for colored visualizations |
max_components |
Maximum number of PCA components to keep (default: 5), or fewer if the data has fewer numeric columns |
k_range |
Range of k values for clustering (default: 2:6). Each is a whole number from 2 to one less than the number of rows, the range a silhouette is defined over. |
Value
A list with class "tidylearn_eda" containing:
- data
The original data frame.
- response
The response variable name, or
NULL.- pca
The fitted PCA model, keeping the first
max_componentscomponents, asprcomp(rank. = max_components)would.- optimal_k
List with optimal cluster count results.
- kmeans
The fitted k-means model.
- hclust
The fitted hierarchical clustering model.
- summary
List with
n_obs,n_vars,n_components(the number kept), andbest_k.
Examples
eda <- tl_explore(iris, response = "Species")
plot(eda)
Extract importance from a tree-based model
Description
Extract importance from a tree-based model
Usage
tl_extract_importance(model)
Arguments
model |
A tidylearn model object |
Value
A data frame with feature importance values, rescaled so the largest is 100. Empty for a tree with no splits.
Fit a gradient boosting model
Description
Fit a gradient boosting model
Usage
tl_fit_boost(
data,
formula,
is_classification = FALSE,
n.trees = 100,
interaction.depth = 3,
shrinkage = 0.1,
n.minobsinnode = 10,
cv.folds = 0,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
n.trees |
Number of trees (default: 100) |
interaction.depth |
Depth of interactions (default: 3) |
shrinkage |
Learning rate (default: 0.1) |
n.minobsinnode |
Minimum number of observations in terminal nodes (default: 10) |
cv.folds |
Number of cross-validation folds (default: 0, no CV) |
... |
Additional arguments to pass to gbm(), including case
|
Details
The distribution follows the response: "gaussian" for
regression, "bernoulli" for two classes and
"multinomial" for more. gbm describes its multinomial
distribution as currently broken, kept only for backwards
compatibility, and warns to that effect on every multiclass fit. For
three or more classes, method = "forest" or
method = "xgboost" are the better supported choices.
Value
A fitted gradient boosting model
Fit a deep learning model
Description
Fit a deep learning model
Usage
tl_fit_deep(
data,
formula,
is_classification = FALSE,
hidden_layers = c(32, 16),
activation = "relu",
dropout = 0.2,
epochs = 30,
batch_size = 32,
validation_split = 0.2,
learning_rate = NULL,
verbose = 0,
...,
compute = "cpu"
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
|
Vector of units in each hidden layer (default: c(32, 16)) | |
activation |
Activation function for hidden layers (default: "relu") |
dropout |
Dropout rate for regularization (default: 0.2) |
epochs |
Number of training epochs (default: 30) |
batch_size |
Batch size for training (default: 32) |
validation_split |
Proportion of the rows held out to validate on,
drawn at random (default: 0.2). Their row numbers are kept as
|
learning_rate |
Optimizer learning rate. NULL (default) leaves keras's own adam default in place. |
verbose |
Verbosity mode (0 = silent, 1 = progress bar, 2 = one line per epoch) (default: 0) |
... |
Additional arguments to pass to keras's fit(). Case
|
compute |
Compute tier. Either |
Value
A fitted deep learning model
Fit an Elastic Net regression model
Description
Fit an Elastic Net regression model
Usage
tl_fit_elastic_net(
data,
formula,
is_classification = FALSE,
alpha = 0.5,
lambda = NULL,
cv_folds = 5,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
alpha |
Mixing parameter (default: 0.5 for Elastic Net) |
lambda |
Regularization parameter: a single penalty, or NULL or a sequence of penalties for cross-validation to choose from |
cv_folds |
Number of folds for cross-validation (default: 5) |
... |
Additional arguments to pass to glmnet() or cv.glmnet() |
Value
A fitted Elastic Net regression model
Fit a random forest model
Description
Fit a random forest model
Usage
tl_fit_forest(
data,
formula,
is_classification = FALSE,
ntree = 500,
mtry = NULL,
importance = TRUE,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
ntree |
Number of trees to grow (default: 500) |
mtry |
Number of variables randomly sampled at each split. Left
to |
importance |
Whether to compute variable importance (default: TRUE) |
... |
Additional arguments to pass to randomForest(). An offset is refused: randomForest leaves it out of the fit. |
Value
A fitted random forest model
Fit a Lasso regression model
Description
Fit a Lasso regression model
Usage
tl_fit_lasso(
data,
formula,
is_classification = FALSE,
alpha = 1,
lambda = NULL,
cv_folds = 5,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
alpha |
Mixing parameter (0 for Ridge, 1 for Lasso, between 0-1 for Elastic Net) |
lambda |
Regularization parameter: a single penalty, or NULL or a sequence of penalties for cross-validation to choose from |
cv_folds |
Number of folds for cross-validation (default: 5) |
... |
Additional arguments to pass to glmnet() or cv.glmnet() |
Value
A fitted Lasso regression model
Fit a linear regression model
Description
Fit a linear regression model
Usage
tl_fit_linear(data, formula, ...)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
... |
Additional arguments to pass to lm() |
Value
A fitted linear regression model
Fit a logistic regression model
Description
Fit a logistic regression model
Usage
tl_fit_logistic(data, formula, ...)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
... |
Additional arguments to pass to glm() |
Value
A fitted logistic regression model
Fit a neural network model
Description
Fit a neural network model
Usage
tl_fit_nn(
data,
formula,
is_classification = FALSE,
size = 5,
decay = 0,
maxit = 100,
trace = FALSE,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
size |
Number of units in the hidden layer (default: 5) |
decay |
Weight decay parameter (default: 0) |
maxit |
Maximum number of iterations (default: 100) |
trace |
Logical; whether to print progress (default: FALSE) |
... |
Additional arguments to pass to nnet(), including case
|
Value
A fitted neural network model
Fit a polynomial regression model
Description
Fit a polynomial regression model
Usage
tl_fit_polynomial(data, formula, degree = 2, ...)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
degree |
Degree of the polynomial (default: 2) |
... |
Additional arguments to pass to lm() |
Value
A fitted polynomial regression model
Fit a regularized regression model
Description
Fits Ridge, Lasso, or Elastic Net regularization.
Usage
tl_fit_regularized(
data,
formula,
is_classification = FALSE,
alpha = 0,
lambda = NULL,
cv_folds = 5,
...,
weights = NULL,
foldid = NULL,
subset = NULL,
offset = NULL
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
alpha |
Mixing parameter (0 for Ridge, 1 for Lasso, between 0-1 for Elastic Net) |
lambda |
Regularization parameter: a single penalty to fit at, or
|
cv_folds |
Number of folds for cross-validation (default: 5) |
... |
Additional arguments to pass to glmnet() or cv.glmnet(). A
name neither function takes is an error, as are |
weights |
Optional case weights, one per row of |
foldid |
Optional fold for each row of |
subset |
Optional rows of |
offset |
Not supported: glmnet would need the offset again at every prediction. An error if supplied. |
Value
A fitted regularized regression model
Fit a Ridge regression model
Description
Fit a Ridge regression model
Usage
tl_fit_ridge(
data,
formula,
is_classification = FALSE,
alpha = 0,
lambda = NULL,
cv_folds = 5,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
alpha |
Mixing parameter (0 for Ridge, 1 for Lasso, between 0-1 for Elastic Net) |
lambda |
Regularization parameter: a single penalty, or NULL or a sequence of penalties for cross-validation to choose from |
cv_folds |
Number of folds for cross-validation (default: 5) |
... |
Additional arguments to pass to glmnet() or cv.glmnet() |
Value
A fitted Ridge regression model
Fit a support vector machine model
Description
Fit a support vector machine model
Usage
tl_fit_svm(
data,
formula,
is_classification = FALSE,
kernel = "radial",
cost = 1,
gamma = NULL,
degree = 3,
tune = FALSE,
tune_folds = 5,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
kernel |
Kernel function ("linear", "polynomial", "radial", "sigmoid") |
cost |
Cost parameter (default: 1) |
gamma |
Gamma parameter for kernels. Left to |
degree |
Degree for polynomial kernel (default: 3) |
tune |
Logical indicating whether to tune hyperparameters (default: FALSE) |
tune_folds |
Number of folds for cross-validation during tuning (default: 5) |
... |
Additional arguments to pass to svm(). |
Value
A fitted SVM model
Fit a decision tree model
Description
Fit a decision tree model
Usage
tl_fit_tree(
data,
formula,
is_classification = FALSE,
cp = 0.01,
minsplit = 20,
maxdepth = 30,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
cp |
Complexity parameter (default: 0.01) |
minsplit |
Minimum number of observations in a node for a split |
maxdepth |
Maximum depth of the tree |
... |
Additional arguments to pass to rpart() or rpart.control(). Any other name is refused, as is an offset, which rpart's predict() does not apply. |
Value
A fitted decision tree model
Fit an XGBoost model
Description
Fit an XGBoost model
Usage
tl_fit_xgboost(
data,
formula,
is_classification = FALSE,
nrounds = 100,
max_depth = 6,
eta = 0.3,
subsample = 1,
colsample_bytree = 1,
min_child_weight = 1,
gamma = 0,
alpha = 0,
lambda = 1,
early_stopping_rounds = NULL,
nthread = NULL,
verbose = 0,
...,
compute = "cpu"
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a classification problem |
nrounds |
Number of boosting rounds (default: 100) |
max_depth |
Maximum depth of trees (default: 6) |
eta |
Learning rate (default: 0.3) |
subsample |
Subsample ratio of observations (default: 1) |
colsample_bytree |
Subsample ratio of columns (default: 1) |
min_child_weight |
Minimum sum of instance weight needed in a child (default: 1) |
gamma |
Minimum loss reduction to make a further partition (default: 0) |
alpha |
L1 regularization term (default: 0) |
lambda |
L2 regularization term (default: 1) |
early_stopping_rounds |
Early stopping rounds (default: NULL). It
needs data to stop on, which a fit on all the rows does not have, so
it is refused unless a validation set is passed as |
nthread |
Number of threads (default: max available) |
verbose |
Verbose output (default: 0) |
... |
Arguments |
compute |
Compute tier. Either |
Value
A fitted XGBoost model
Get the best model from a pipeline
Description
Get the best model from a pipeline
Usage
tl_get_best_model(pipeline)
Arguments
pipeline |
A tidylearn pipeline object with results |
Value
The best tidylearn_model object from the pipeline,
selected by the metric specified in evaluation$best_metric.
The model was fitted on the preprocessed training data, so under the
default preprocessing it expects predictors imputed and standardised
the same way. Predict through tl_predict_pipeline, which
replays that preprocessing on raw rows; predict() on the model
itself reads raw values as if they were already standardised, and
returns wrong predictions without a warning.
Examples
pipe <- tl_pipeline(iris, Species ~ .,
models = list(tree = list(method = "tree")),
evaluation = list(metrics = "accuracy", validation = "cv",
cv_folds = 2, best_metric = "accuracy"))
pipe <- tl_run_pipeline(pipe, verbose = FALSE)
best <- tl_get_best_model(pipe)
best$spec$method
# Predict through the pipeline, which preprocesses the new rows first
tl_predict_pipeline(pipe, iris[c(1, 51, 101), ])
Extract importance from a regularized regression model
Description
Extract importance from a regularized regression model
Usage
tl_get_importance_regularized(model, lambda = "1se")
Arguments
model |
A tidylearn regularized model object |
lambda |
Which lambda to use: "1se" (default), "min", or a numeric penalty within the fitted path |
Value
A data frame with feature importance values: each coefficient's absolute value times its predictor's standard deviation, so the ranking does not depend on units, rescaled to a maximum of 100. For a multiclass model a predictor takes its largest value across classes.
Calculate influence measures for a linear model
Description
Calculate influence measures for a linear model
Usage
tl_influence_measures(
model,
threshold_cook = NULL,
threshold_leverage = NULL,
threshold_dffits = NULL
)
Arguments
model |
A tidylearn model object |
threshold_cook |
Cook's distance threshold (default: 4/n) |
threshold_leverage |
Leverage threshold (default: 2*(p+1)/n) |
threshold_dffits |
DFFITS threshold (default: 2*sqrt((p+1)/n)) |
Value
A data frame with one row per observation containing influence
measures: cooks_distance, leverage, dffits,
std_residual, stud_residual, boolean flags for each
threshold (is_cook_influential, is_leverage_influential,
is_dffits_influential, is_outlier), per-coefficient
dfbetas_* columns, and an overall is_influential flag.
Threshold values are stored as attributes.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_influence_measures(model)
Calculate partial effects based on a model with interactions
Description
Calculate partial effects based on a model with interactions
Usage
tl_interaction_effects(model, var, by_var, at_values = NULL, intervals = TRUE)
Arguments
model |
A tidylearn model object |
var |
Variable to calculate effects for |
by_var |
Variable to calculate effects by (interaction variable) |
at_values |
Named list of values at which to hold other variables. A variable not named is held at its median if it is numeric, and otherwise at its most frequent value, a tie going to the earlier level. |
intervals |
Logical; whether to add 95\
|
Value
For numeric var: a list with effects (data frame of
predicted values across the variable range for each value of
by_var) and slopes (data frame with the slope of
var at each value of by_var). For categorical
var: a data frame of predicted values at each factor level for
each level of by_var. A numeric by_var is evaluated at its
quartiles; quartiles that tie are evaluated once, with a label naming
each quartile they stand for, such as "Q0/Q25". A factor,
character or logical variable is categorical, and is evaluated at the
values the data holds: a level no row uses, and a missing value, are
left out. A factor is returned as a factor with the model's levels.
fit, lower, upper and slope are on the
response scale whatever intervals is set to. For a
classification model that is the predicted probability of the second
class, so the response must have two classes. slope is the slope
of a straight line fitted to fit across the range of var,
so for a non-linear link it is an average rate of change over that
range. A model variable named fit, by_value or
by_label, or with intervals lower or upper, would
be overwritten by these columns and is refused.
slopes$slope_se is the standard error of a straight line fitted
to the prediction grid, not the sampling uncertainty of the marginal
effect. For a linear model the grid is exactly linear in var,
so this is near zero by construction and should not be read as a
precise estimate. Use summary(model$fit) for inference on the
interaction coefficient itself.
Examples
model <- tl_model(mtcars, mpg ~ wt * hp, method = "linear")
# How the effect of weight changes across horsepower
effects <- tl_interaction_effects(model, var = "wt", by_var = "hp")
head(effects$effects)
effects$slopes
# slopes$slope_se describes the fitted grid, not the sampling
# uncertainty of the marginal effect -- for that, read the coefficient
summary(model$fit)$coefficients["wt:hp", ]
Load a pipeline from disk
Description
Load a pipeline from disk
Usage
tl_load_pipeline(file)
Arguments
file |
Path to the pipeline file |
Value
A tidylearn_pipeline object previously saved with
tl_save_pipeline.
Examples
pipe <- tl_pipeline(iris, Species ~ .)
f <- tempfile(fileext = ".rds")
tl_save_pipeline(pipe, f)
pipe2 <- tl_load_pipeline(f)
Create a tidylearn model
Description
Unified interface for creating machine learning models by wrapping established R packages. This function dispatches to the appropriate underlying package based on the method.
Usage
tl_model(data, formula = NULL, method = "linear", ..., compute = "cpu")
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model. For
unsupervised methods, use |
method |
The modeling method. Supervised: "linear"
(stats::lm), "polynomial" (stats::lm on polynomial terms),
"logistic" (stats::glm), "tree" (rpart),
"forest" (randomForest), "boost" (gbm),
"ridge"/"lasso"/"elastic_net" (glmnet), "svm" (e1071),
"nn" (nnet), "deep" (keras), "xgboost" (xgboost).
The method and the response have to agree, and a mismatch is an
error rather than a meaningless fit: |
... |
Arguments for the method: see the Method arguments section. Anything else is passed to the underlying model function. |
compute |
Compute tier for the fit. One of |
Details
The wrapped packages include: stats (lm, glm, prcomp, kmeans, hclust), glmnet, randomForest, xgboost, gbm, e1071, nnet, rpart, cluster, and dbscan. The underlying algorithms are unchanged - this function provides a consistent interface and returns tidy output.
For a supervised method, model$fit is the object the wrapped
function returned. An unsupervised method returns tidied components as
well, so the wrapped object sits at model$fit$model and
model$fit is the list holding both.
For classification, the response is reduced to the classes it actually contains: subsetting a data frame keeps every factor level, and a level no row uses would otherwise be reported as a class, given its own (zero) probability column, and counted when deciding whether the problem is binary. The fit is unaffected.
Whether a supervised model is a classification or a regression is
decided by the response the formula computes, so factor(cyl) ~ wt
is a classification even though cyl is numeric. A factor or text
response is a classification. A logical response is a regression for
every method but "logistic" – with "linear", a linear
probability model – so write factor(y) ~ ... to classify it.
A categorical predictor stored as text, as tl_read() returns it,
is made a factor before the fit and stored as one in $data, so
every method treats it as a category. At predict(), categorical
columns in new data are read against the levels seen in training: new
data may hold only some of them, and a level the model was not trained
on is an error that names the column.
Value
A tidylearn_model object (S3) containing the fitted model
($fit, or $fit$model for an unsupervised method),
model specification ($spec), and training data
($data). update() and step() on $fit
refit on the training rows and weights, whatever is in the calling
environment. The object also inherits from a method-specific class
(e.g., tidylearn_linear) and a paradigm class
(tidylearn_supervised or tidylearn_unsupervised).
Method arguments
Arguments in ... are passed to the function the method wraps,
except for these, which tidylearn takes itself:
"polynomial"degree(default 2). Each numeric main effect is replaced bypoly(term, degree, raw = TRUE)and the result fitted withlm(). A numeric term is one that computes a numeric vector, such aswtorlog(wt), or a one-column matrix, such asscale(wt). One that is also part of an interaction keeps its own term and gainsI(x^2)up toI(x^degree), so the interaction is coded as written. Factor and other non-numeric terms, interactions,I()terms, bases such aspoly()or a spline's, the response as written, anoffset()and a removed intercept are kept as they are."ridge","lasso","elastic_net"-
alpha, glmnet's mixing parameter (by default 0, 1 and 0.5);lambda, a single penalty to fit at, a sequence of penalties forglmnet::cv.glmnet()to choose from, orNULL(the default) to let it choose its own; andcv_folds(default 5), the number of folds for that cross-validation, which takes the place of glmnet'snfolds.predict()uses thelambda.1sepenalty. tidylearn setsx,y,familyandnfoldsitself and refuses them, along with any argument glmnet does not take. "svm"tune(defaultFALSE) andtune_folds(default 5), to choosecostby cross-validation before the fit, withgammafor a non-linear kernel anddegreefor a polynomial one."deep"hidden_layers,activation,dropout,epochs,batch_size,validation_splitandlearning_rate."pca"scaleandcenter(bothTRUE)."mds"mds_method, the variant:"classical"(the default,stats::cmdscale()),"metric"or"nonmetric"(smacof), or"sammon"or"kruskal"(MASS).k, or its aliasndim, is the number of dimensions (default 2)."kmeans","pam","clara"k, the number of clusters (default 3); for"pam",metricas well (default"euclidean")."hclust"hclust_method, the linkage:"average"(the default),"ward.D","ward.D2","single","complete","mcquitty","median"or"centroid"; anddistance(default"euclidean")."dbscan"eps(default 0.5),minPts(default 5) anddistance(default"euclidean").
Some defaults differ from the wrapped function's: "forest" computes
importance (importance = TRUE), "boost" grows trees of
interaction.depth = 3, and "nn" fits size = 5
hidden units with trace = FALSE.
weights and subset take values, one per row of
data (such as weights = data$w), not column names. A
subset is applied before the fit, and $data holds only
the rows it selects. Case weights are applied by every supervised method
except "svm" and "deep", which refuse them. An offset,
written as offset() in the formula, is applied by
"linear", "polynomial" and "logistic"; the other
methods refuse one, because their predict() would not add it
back.
Examples
# Classification -> wraps randomForest::randomForest()
model <- tl_model(iris, Species ~ ., method = "forest")
model$fit # Access the raw randomForest object
# Regression -> wraps stats::lm()
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
model$fit # Access the raw lm object
# PCA -> wraps stats::prcomp()
model <- tl_model(iris, ~ ., method = "pca")
model$fit$model # The raw prcomp object, alongside tidied components
# Clustering -> wraps stats::kmeans()
model <- tl_model(iris, method = "kmeans", k = 3)
model$fit$model # The raw kmeans object
Create a modeling pipeline
Description
Create a modeling pipeline
Usage
tl_pipeline(
data,
formula,
preprocessing = NULL,
models = NULL,
evaluation = NULL,
...
)
Arguments
data |
A data frame containing the data |
formula |
A formula specifying the model |
preprocessing |
A named list of preprocessing switches, each
|
models |
A list of models to train |
evaluation |
A list of evaluation criteria |
... |
Not used. Anything passed here is an error, so a misspelt
argument such as |
Value
A tidylearn_pipeline object (S3 list) with components
$formula, $data, $preprocessing,
$models, $evaluation, and $results
(initially NULL; populated after tl_run_pipeline).
Examples
pipe <- tl_pipeline(iris, Species ~ .,
models = list(tree = list(method = "tree")))
print(pipe)
Plot actual vs predicted values for a regression model
Description
Plot actual vs predicted values for a regression model
Usage
tl_plot_actual_predicted(model, new_data = NULL, ...)
Arguments
model |
A tidylearn regression model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
... |
Additional arguments |
Value
A ggplot object
Plot calibration curve for a classification model
Description
Plot calibration curve for a classification model
Usage
tl_plot_calibration(model, new_data = NULL, bins = 10, ...)
Arguments
model |
A tidylearn classification model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
bins |
Number of bins for grouping predictions (default: 10) |
... |
Additional arguments |
Value
A ggplot object with calibration curve
Plot confusion matrix for a classification model
Description
Plot confusion matrix for a classification model
Usage
tl_plot_confusion(model, new_data = NULL, ...)
Arguments
model |
A tidylearn classification model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
... |
Additional arguments |
Value
A ggplot object with confusion matrix
Plot comparison of cross-validation results
Description
Plot comparison of cross-validation results
Usage
tl_plot_cv_comparison(cv_results, metrics = NULL)
Arguments
cv_results |
Results from tl_compare_cv function |
metrics |
Character vector of metrics to plot (if NULL, plots all metrics) |
Value
A ggplot object showing boxplots of
cross-validation metric distributions for each model.
Examples
m1 <- tl_model(mtcars, mpg ~ wt, method = "linear")
m2 <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
cv <- tl_compare_cv(mtcars, list(simple = m1, full = m2), folds = 3)
tl_plot_cv_comparison(cv)
Plot cross-validation results
Description
Plot cross-validation results
Usage
tl_plot_cv_results(cv_results, metrics = NULL)
Arguments
cv_results |
Cross-validation results from tl_cv function |
metrics |
Character vector of metrics to plot (if NULL, plots all metrics) |
Value
A ggplot object.
Examples
cv <- tl_cv(mtcars, mpg ~ wt + hp, method = "linear", folds = 5)
tl_plot_cv_results(cv)
# One metric rather than every one the folds scored
tl_plot_cv_results(cv, metrics = "rmse")
Plot deep learning model architecture
Description
Plot deep learning model architecture
Usage
tl_plot_deep_architecture(model, ...)
Arguments
model |
A tidylearn deep learning model object |
... |
Additional arguments passed to the |
Value
NULL, invisibly. Called for its side effect: keras draws
the architecture diagram on the current graphics device, or writes it
to to_file. keras renders it through the Python packages
pydot and graphviz, and errors saying so when they are
not installed.
Examples
## Not run:
if (requireNamespace("keras", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "deep", epochs = 5)
tl_plot_deep_architecture(model)
}
## End(Not run)
Plot deep learning model training history
Description
Plot deep learning model training history
Usage
tl_plot_deep_history(model, metrics = c("loss", "val_loss"), ...)
Arguments
model |
A tidylearn deep learning model object |
metrics |
Which metrics to plot (default: c("loss", "val_loss")) |
... |
Additional arguments |
Value
A ggplot object.
Examples
## Not run:
if (requireNamespace("keras", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "deep", epochs = 5)
tl_plot_deep_history(model)
}
## End(Not run)
Plot diagnostics for a regression model
Description
Plot diagnostics for a regression model
Usage
tl_plot_diagnostics(model, which = 1:4, ...)
Arguments
model |
A tidylearn regression model object |
which |
Which plots to create (1:4) |
... |
Additional arguments |
Value
A ggplot object (or list of ggplot objects)
Plot gain chart for a classification model
Description
Plot gain chart for a classification model
Usage
tl_plot_gain(model, new_data = NULL, bins = 10, ...)
Arguments
model |
A tidylearn classification model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
bins |
Number of bins for grouping predictions (default: 10) |
... |
Additional arguments |
Value
A ggplot object.
Examples
iris_bin <- iris[iris$Species != "setosa", ]
iris_bin$Species <- factor(iris_bin$Species)
model <- tl_model(iris_bin, Species ~ ., method = "logistic")
tl_plot_gain(model)
Plot variable importance for tree-based models
Description
Plot variable importance for tree-based models
Usage
tl_plot_importance(model, top_n = 20, ...)
Arguments
model |
A tidylearn tree-based model object |
top_n |
Number of top features to display (default: 20) |
... |
Additional arguments |
Value
A ggplot object
Plot feature importance across multiple models
Description
Each model's importance is rescaled on its own: its largest value
becomes 100, or, when no value is positive (a forest's permutation
importance can be negative throughout), its largest magnitude becomes
-100. A factor predictor
appears once, under its own name, for every model: the largest of its
design columns stands for it where a method ranks those columns
separately (ridge, lasso, elastic net and xgboost). A predictor a model
was given but did not use scores zero for that model; one it was never
given has no bar for it. Tree, forest and boost models are given the
variables of an interaction such as wt:hp rather than the
interaction itself, so it has no bar for them. Features are ranked on
their mean importance over the models that were given them.
Usage
tl_plot_importance_comparison(..., top_n = 10, names = NULL)
Arguments
... |
tidylearn model objects to compare |
top_n |
Number of top features to display (default: 10) |
names |
Optional character vector of model names, one unique name per model |
Value
A ggplot object.
Examples
m1 <- tl_model(iris, Sepal.Length ~ ., method = "forest")
m2 <- tl_model(iris, Sepal.Length ~ ., method = "boost")
tl_plot_importance_comparison(m1, m2, names = c("Forest", "Boost"))
Plot variable importance for a regularized model
Description
Plot variable importance for a regularized model
Usage
tl_plot_importance_regularized(model, lambda = "1se", top_n = 20, ...)
Arguments
model |
A tidylearn regularized model object |
lambda |
Which lambda to use: "1se" (default), "min", or a numeric penalty within the fitted path |
top_n |
Number of top features to display (default: 20) |
... |
Additional arguments |
Value
A ggplot object.
Examples
model <- tl_model(mtcars, mpg ~ ., method = "lasso")
tl_plot_importance_regularized(model)
Plot influence diagnostics
Description
Plot influence diagnostics
Usage
tl_plot_influence(
model,
plot_type = "cook",
threshold_cook = NULL,
threshold_leverage = NULL,
threshold_dffits = NULL,
n_labels = 3,
label_size = 3
)
Arguments
model |
A tidylearn model object |
plot_type |
Type of influence plot: "cook", "leverage", "index" |
threshold_cook |
Cook's distance threshold (default: 4/n) |
threshold_leverage |
Leverage threshold (default: 2*(p+1)/n) |
threshold_dffits |
DFFITS threshold (default: 2*sqrt((p+1)/n)) |
n_labels |
Number of points to label (default: 3) |
label_size |
Text size for labels (default: 3) |
Value
A ggplot object.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_plot_influence(model, plot_type = "cook")
Plot interaction effects
Description
Plot interaction effects
Usage
tl_plot_interaction(
model,
var1,
var2,
n_points = 100,
fixed_values = NULL,
confidence = TRUE,
...
)
Arguments
model |
A tidylearn model object |
var1 |
First variable in the interaction |
var2 |
Second variable in the interaction |
n_points |
Number of points to use for continuous variables |
fixed_values |
Named list of values for other variables in the model. A variable not named is held at its median if it is numeric, and otherwise at its most frequent value, a tie going to the earlier level. |
confidence |
Logical; whether to show a 95\
band is drawn when one variable is numeric and the other categorical,
and needs a model whose underlying fit is an |
... |
Additional arguments to pass to predict() |
Value
A ggplot object. Two numeric variables are
drawn as a filled contour of the prediction; a numeric and a categorical
variable as one line per category; two categorical variables as dodged
bars. A factor, character or logical variable is categorical, and only
the values the data holds are drawn. For a classification model the
prediction is the probability of the second class, so the response
must have two classes. The plot's data holds the prediction in a column
named prediction, and a band in .lower and .upper,
so a model variable with one of those names is refused.
Examples
model <- tl_model(mtcars, mpg ~ wt * hp, method = "linear")
# Two numeric variables are drawn as a filled contour over both ranges
tl_plot_interaction(model, var1 = "wt", var2 = "hp")
# A numeric by categorical interaction is drawn as one line per level,
# each with a confidence band
am_model <- tl_model(transform(mtcars, am = factor(am)), mpg ~ wt * am,
method = "linear")
tl_plot_interaction(am_model, var1 = "wt", var2 = "am")
# Coarser grid, no band
tl_plot_interaction(am_model, var1 = "wt", var2 = "am",
n_points = 20, confidence = FALSE)
Create confidence and prediction interval plots
Description
Create confidence and prediction interval plots
Usage
tl_plot_intervals(model, new_data = NULL, level = 0.95, ...)
Arguments
model |
A tidylearn regression model object |
new_data |
Optional data frame for prediction (if NULL, uses training data) |
level |
Confidence level (default: 0.95) |
... |
Additional arguments |
Value
A ggplot object.
Examples
model <- tl_model(mtcars, mpg ~ wt, method = "linear")
tl_plot_intervals(model)
Plot lift chart for a classification model
Description
Plot lift chart for a classification model
Usage
tl_plot_lift(model, new_data = NULL, bins = 10, ...)
Arguments
model |
A tidylearn classification model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
bins |
Number of bins for grouping predictions (default: 10) |
... |
Additional arguments |
Value
A ggplot object.
Examples
iris_bin <- iris[iris$Species != "setosa", ]
iris_bin$Species <- factor(iris_bin$Species)
model <- tl_model(iris_bin, Species ~ ., method = "logistic")
tl_plot_lift(model)
Plot a supervised tidylearn model
Description
Dispatches to the appropriate plotting function based on model type and requested plot type.
Usage
tl_plot_model(model, type = "auto", ...)
Arguments
model |
A tidylearn supervised model object |
type |
Plot type. For regression: "auto", "actual_predicted",
"residuals", "diagnostics". For classification: "auto", "confusion",
"roc", "precision_recall", "calibration", "lift", "gain".
"importance" is available for tree-based and regularized models.
"diagnostics" needs a model fitted by |
... |
Additional arguments passed to the underlying plot function |
Value
A ggplot2 object (invisibly for base-graphics plots)
Plot model comparison
Description
Plot model comparison
Usage
tl_plot_model_comparison(..., new_data = NULL, metrics = NULL, names = NULL)
Arguments
... |
tidylearn model objects to compare |
new_data |
Optional data frame for evaluation. If NULL, the models
are scored on their training data, which they must share: models
fitted on different data are an error asking for |
metrics |
Character vector of metrics to compute |
names |
Optional character vector of model names |
Value
A ggplot object.
Examples
m1 <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
m2 <- tl_model(mtcars, mpg ~ wt + hp, method = "lasso")
tl_plot_model_comparison(m1, m2, names = c("Linear", "Lasso"))
Plot neural network architecture
Description
Plot neural network architecture
Usage
tl_plot_nn_architecture(model, ...)
Arguments
model |
A tidylearn neural network model object |
... |
Additional arguments |
Value
The return value of plotnet, called for
its side effect of drawing the network diagram, or NULL if the
NeuralNetTools package is not installed.
Examples
if (requireNamespace("NeuralNetTools", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "nn", size = 3)
tl_plot_nn_architecture(model)
}
Plot a neural network tuning grid
Description
Draws the size-by-decay grid as a heatmap of cross-validated error.
Usage
tl_plot_nn_tuning(model, ...)
Arguments
model |
The list returned by |
... |
Additional arguments |
Value
A ggplot object.
Examples
tuned <- tl_tune_nn(iris, Species ~ .,
is_classification = TRUE,
sizes = c(2, 5), decays = c(0, 0.01), folds = 3)
# The tuning result itself, not tuned$model
tl_plot_nn_tuning(tuned)
Plot partial dependence for tree-based models
Description
Plot partial dependence for tree-based models
Usage
tl_plot_partial_dependence(model, var, n.pts = 20, ...)
Arguments
model |
A tidylearn tree-based model object |
var |
Variable name to plot |
n.pts |
Number of points for continuous variables (default: 20) |
... |
Additional arguments |
Value
A ggplot object. Its data has a
var_value column and the mean prediction over the model's
training rows, y, at each value. For classification, y
is a mean class probability and a class column says which
class: the positive class (the second level) alone for a two-class
model, and every class, one line each, for more.
Examples
model <- tl_model(mtcars, mpg ~ ., method = "forest")
tl_plot_partial_dependence(model, var = "wt")
Plot precision-recall curve for a classification model
Description
Plot precision-recall curve for a classification model
Usage
tl_plot_precision_recall(model, new_data = NULL, ...)
Arguments
model |
A tidylearn classification model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
... |
Additional arguments |
Value
A ggplot object with precision-recall curve
Plot cross-validation results for a regularized model
Description
Shows the cross-validation error as a function of lambda for ridge, lasso, or elastic net models fitted with cv.glmnet.
Usage
tl_plot_regularization_cv(model, ...)
Arguments
model |
A tidylearn regularized model object (ridge, lasso, or elastic_net) |
... |
Additional arguments (currently unused) |
Value
A ggplot object.
Examples
model <- tl_model(mtcars, mpg ~ ., method = "ridge")
tl_plot_regularization_cv(model)
Plot regularization path for a regularized model
Description
Plot regularization path for a regularized model
Usage
tl_plot_regularization_path(model, label_n = 5, ...)
Arguments
model |
A tidylearn regularized model object |
label_n |
Number of top features to label (default: 5) |
... |
Additional arguments |
Value
A ggplot object.
Examples
model <- tl_model(mtcars, mpg ~ ., method = "lasso")
tl_plot_regularization_path(model)
Plot residuals for a regression model
Description
Plot residuals for a regression model
Usage
tl_plot_residuals(model, type = "fitted", ...)
Arguments
model |
A tidylearn regression model object |
type |
Type of residual plot: "fitted" (default), "histogram", "predicted" |
... |
Additional arguments |
Value
A ggplot object
Plot ROC curve for a classification model
Description
Plot ROC curve for a classification model
Usage
tl_plot_roc(model, new_data = NULL, ...)
Arguments
model |
A tidylearn classification model object |
new_data |
Optional data frame for evaluation (if NULL, uses training data) |
... |
Additional arguments |
Value
A ggplot object with ROC curve
Plot SVM decision boundary
Description
Plot SVM decision boundary
Usage
tl_plot_svm_boundary(model, x_var = NULL, y_var = NULL, grid_size = 100, ...)
Arguments
model |
A tidylearn SVM model object |
x_var |
Name of the x-axis variable. Defaults to the first numeric predictor in the model's formula. |
y_var |
Name of the y-axis variable. Defaults to the next numeric predictor in the model's formula. |
grid_size |
Number of points in each dimension for the grid (default: 100) |
... |
Additional arguments |
Value
A ggplot object. The other predictors are
held at their mean, or their most frequent level. A two-class model
fitted with probabilities also gets the 0.5 probability contour.
Examples
if (requireNamespace("e1071", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "svm")
tl_plot_svm_boundary(model,
x_var = "Sepal.Length", y_var = "Sepal.Width")
}
Plot SVM tuning results
Description
Plot SVM tuning results
Usage
tl_plot_svm_tuning(model, ...)
Arguments
model |
A tidylearn SVM model object |
... |
Additional arguments |
Value
A ggplot object.
Examples
if (requireNamespace("e1071", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "svm",
kernel = "linear", tune = TRUE, tune_folds = 2)
tl_plot_svm_tuning(model)
}
Plot a decision tree
Description
Plot a decision tree
Usage
tl_plot_tree(model, ...)
Arguments
model |
A tidylearn tree model object |
... |
Additional arguments to pass to rpart.plot() |
Value
The return value of rpart.plot, called
for its side effect of drawing the tree.
Examples
model <- tl_model(iris, Species ~ ., method = "tree")
tl_plot_tree(model)
Plot hyperparameter tuning results
Description
Plot hyperparameter tuning results
Usage
tl_plot_tuning_results(
model,
top_n = 5,
param1 = NULL,
param2 = NULL,
plot_type = "scatter"
)
Arguments
model |
A tidylearn model object with tuning results |
top_n |
Number of top parameter sets to highlight |
param1 |
First parameter to plot (for 2D grid or scatter plots) |
param2 |
Second parameter to plot (for 2D grid or scatter plots) |
plot_type |
Type of plot: "scatter", "grid", "parallel", "importance" |
Details
A parameter whose candidates are not single values, such as
hidden_layers = list(c(10), c(20, 10)) or a parms list,
is drawn as a categorical one, each value labelled as the verbose
messages print it. The importance of a numeric parameter is the
absolute correlation of its values with the score; that of a
categorical one is eta squared from a one-way ANOVA of the score,
and 0 when the sets that were scored all share one value.
Value
A ggplot object.
Examples
model <- tl_tune_grid(iris, Species ~ ., method = "tree",
param_grid = list(cp = c(0.01, 0.1), minsplit = c(10, 20)),
folds = 2, verbose = FALSE)
tl_plot_tuning_results(model)
Plot an unsupervised tidylearn model
Description
Dispatches to the appropriate plotting function based on the unsupervised model method.
Usage
tl_plot_unsupervised(model, type = "auto", ...)
Arguments
model |
A tidylearn unsupervised model object |
type |
Plot type (default: "auto"). Currently unused; reserved for future sub-type selection. |
... |
Additional arguments passed to the underlying plot function |
Value
A ggplot2 object or invisible result
Plot feature importance for an XGBoost model
Description
Plot feature importance for an XGBoost model
Usage
tl_plot_xgboost_importance(model, top_n = 10, importance_type = "gain", ...)
Arguments
model |
A tidylearn XGBoost model object |
top_n |
Number of top features to display (default: 10) |
importance_type |
Type of importance: "gain" (default), "cover",
"frequency" or "weight", read from the matching column of
|
... |
Additional arguments passed to |
Value
A ggplot object. Its data holds the
top_n features and their importance, relative to the
most important feature's.
Examples
if (requireNamespace("xgboost", quietly = TRUE)) {
model <- tl_model(mtcars, mpg ~ ., method = "xgboost", nthread = 2)
tl_plot_xgboost_importance(model)
}
Plot SHAP dependence for a specific feature
Description
Plot SHAP dependence for a specific feature
Usage
tl_plot_xgboost_shap_dependence(
model,
feature,
interaction_feature = NULL,
data = NULL,
n_samples = 100
)
Arguments
model |
A tidylearn XGBoost model object |
feature |
Feature name to plot |
interaction_feature |
Feature to use for coloring (default: NULL) |
data |
Data for SHAP value calculation (default: NULL, uses training data) |
n_samples |
Number of samples to use (default: 100, NULL for all) |
Value
A ggplot object. A multiclass model is
drawn one panel per class, from that class's SHAP values.
Examples
if (requireNamespace("xgboost", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "xgboost", nrounds = 10,
nthread = 2)
# One panel per class
tl_plot_xgboost_shap_dependence(model, feature = "Petal.Length")
# Colour the points by a second feature to read the interaction
tl_plot_xgboost_shap_dependence(model,
feature = "Petal.Length",
interaction_feature = "Petal.Width")
}
Plot SHAP summary for XGBoost model
Description
Plot SHAP summary for XGBoost model
Usage
tl_plot_xgboost_shap_summary(model, data = NULL, top_n = 10, n_samples = 100)
Arguments
model |
A tidylearn XGBoost model object |
data |
Data for SHAP value calculation (default: NULL, uses training data) |
top_n |
Number of top features to display (default: 10) |
n_samples |
Number of samples to use (default: 100, NULL for all) |
Value
A ggplot object. Features are ranked by
mean absolute SHAP value; a multiclass model is drawn one panel per
class.
Examples
if (requireNamespace("xgboost", quietly = TRUE)) {
model <- tl_model(mtcars, mpg ~ ., method = "xgboost", nthread = 2)
tl_plot_xgboost_shap_summary(model, n_samples = 20)
}
Plot XGBoost tree visualization
Description
Plot XGBoost tree visualization
Usage
tl_plot_xgboost_tree(model, tree_index = 0, ...)
Arguments
model |
A tidylearn XGBoost model object |
tree_index |
Index of the tree to plot, counting the first tree as 0 (default: 0). A multiclass model grows one tree per class in each round. |
... |
Additional arguments passed to |
Value
The return value of xgb.plot.tree, a
tree diagram rendered via the DiagrammeR package.
Examples
if (requireNamespace("xgboost", quietly = TRUE) &&
requireNamespace("DiagrammeR", quietly = TRUE)) {
model <- tl_model(iris, Species ~ ., method = "xgboost", nrounds = 10,
nthread = 2)
# tree_index is zero-based, so this is the first tree
tl_plot_xgboost_tree(model, tree_index = 0)
}
Predict using a gradient boosting model
Description
Predict using a gradient boosting model
Usage
tl_predict_boost(model, new_data, type = "response", n.trees = NULL, ...)
Arguments
model |
A tidylearn boost model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "response" (default), "prob" (for classification) |
n.trees |
Number of trees to use for prediction (if NULL, uses optimal number) |
... |
Additional arguments |
Value
Predictions
Predict using a deep learning model
Description
Predict using a deep learning model
Usage
tl_predict_deep(model, new_data, type = "response", ...)
Arguments
model |
A tidylearn deep learning model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "response" (default), "prob" (for classification), "class" (for classification) |
... |
Additional arguments |
Value
Predictions
Predict using a random forest model
Description
Predict using a random forest model
Usage
tl_predict_forest(model, new_data, type = "response", ...)
Arguments
model |
A tidylearn forest model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "response" (default), "prob" (for classification) |
... |
Additional arguments |
Value
Predictions
Predict using a logistic regression model
Description
Predict using a logistic regression model
Usage
tl_predict_logistic(model, new_data, type = "prob", ...)
Arguments
model |
A tidylearn logistic model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "prob" (default), "class", "response" |
... |
Additional arguments |
Value
Predictions
Predict using a neural network model
Description
Predict using a neural network model
Usage
tl_predict_nn(model, new_data, type = "response", ...)
Arguments
model |
A tidylearn neural network model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "response" (default), "prob" (for classification), "class" (for classification) |
... |
Additional arguments |
Value
Predictions
Make predictions using a pipeline
Description
Make predictions using a pipeline
Usage
tl_predict_pipeline(
pipeline,
new_data,
type = "response",
model_name = NULL,
...
)
Arguments
pipeline |
A tidylearn pipeline object with results |
new_data |
A data frame containing the new data |
type |
Type of prediction (default: "response") |
model_name |
Name of model to use (if NULL, uses the best model) |
... |
Additional arguments passed to predict |
Value
A tibble with a .pred column containing
predictions from the selected (or best) pipeline model, after
applying the same preprocessing steps used during training.
Examples
train <- iris[c(1:40, 51:90, 101:140), ]
test <- iris[c(41:50, 91:100, 141:150), ]
pipe <- tl_pipeline(train, Species ~ .,
models = list(
tree = list(method = "tree"),
forest = list(method = "forest", ntree = 100)
),
evaluation = list(validation = "cv", cv_folds = 3))
pipe <- tl_run_pipeline(pipe, verbose = FALSE)
# The best model, with the preprocessing learned on the training rows
tl_predict_pipeline(pipe, test)
# Or a named candidate instead of the winner
tl_predict_pipeline(pipe, test, model_name = "tree")
Predict using a support vector machine model
Description
Predict using a support vector machine model
Usage
tl_predict_svm(model, new_data, type = "response", ...)
Arguments
model |
A tidylearn SVM model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "response" (default), "prob" (for classification) |
... |
Additional arguments |
Value
Predictions
Predict using a decision tree model
Description
Predict using a decision tree model
Usage
tl_predict_tree(model, new_data, type = "response", ...)
Arguments
model |
A tidylearn tree model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "response" (default), "prob" or "class" (for classification) |
... |
Additional arguments |
Value
Predictions
Predict using an XGBoost model
Description
Predict using an XGBoost model
Usage
tl_predict_xgboost(
model,
new_data,
type = "response",
iterationrange = NULL,
ntreelimit = NULL,
...
)
Arguments
model |
A tidylearn XGBoost model object |
new_data |
A data frame containing the new data |
type |
Type of prediction: "response" (default), "prob" (for classification), "class" (for classification) |
iterationrange |
Boosting iterations to predict with, as
|
ntreelimit |
Deprecated. Use |
... |
Additional arguments |
Value
Predictions
Data Preprocessing for tidylearn
Description
Unified preprocessing functions that work with both supervised and unsupervised workflows Prepare Data for Machine Learning
Usage
tl_prepare_data(
data,
formula = NULL,
impute_method = "mean",
scale_method = "standardize",
encode_categorical = TRUE,
remove_zero_variance = TRUE,
remove_correlated = FALSE,
correlation_cutoff = 0.95
)
Arguments
data |
A data frame. A grouped tibble is prepared as a whole and returned ungrouped. |
formula |
Optional two-sided formula (for supervised learning),
whose response must be a column of |
impute_method |
Method for imputing a missing numeric value: "mean", "median" or "mode". A missing categorical value is always filled with the column's most frequent value. |
scale_method |
Scaling method: "standardize",
"normalize", "robust", "none". A column whose spread is zero or not
finite, such as one holding an |
encode_categorical |
Whether to encode categorical variables (default: TRUE) |
remove_zero_variance |
Remove zero-variance features (default: TRUE) |
remove_correlated |
Remove highly correlated features (default: FALSE) |
correlation_cutoff |
Correlation threshold for removal, greater than 0 and at most 1 (default: 0.95) |
Details
Comprehensive preprocessing pipeline including imputation, scaling, encoding, and feature engineering
The statistics are learned from, and applied to, the data passed in.
Preparing a whole dataset and then splitting it lets the test rows
shape the imputation values and scaling their own scores are measured
against. To evaluate a model, split first, or use
tl_pipeline, which learns its preprocessing inside each
resampling fold.
Value
A list with components:
dataThe processed data frame.
original_dataThe original unprocessed data frame.
preprocessing_stepsA record of each step applied (imputation values, encoding maps, scaling parameters, etc.). It is for inspection: no function applies it to new data.
formulaThe formula passed in (or
NULL).
Examples
processed <- tl_prepare_data(iris, Species ~ ., scale_method = "standardize")
model <- tl_model(processed$data, Species ~ ., method = "tree")
Read data from diverse sources
Description
Auto-detects the data format from the file extension or source pattern and
dispatches to the appropriate reader. All readers
return a tidylearn_data object, which is a
tibble subclass carrying metadata about the data
source.
Usage
tl_read(source, ..., format = NULL, .quiet = FALSE)
Arguments
source |
A file path or |
... |
Additional arguments passed to the format-specific reader. |
format |
Optional explicit format override.
One of |
.quiet |
Logical. If |
Details
When source is a character vector of multiple paths, each file is read
and row-bound into a single result with a source_file column
giving each file's path below the deepest folder they share. When
source is a directory path, it is equivalent to calling
tl_read_dir(). When source is a local .zip file, it
is equivalent to calling tl_read_zip().
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Examples
# The format is detected from the extension
csv <- tempfile(fileext = ".csv")
write.csv(mtcars, csv, row.names = FALSE)
tl_read(csv)
# Several files are row-bound, with a source_file column naming each
jan <- tempfile(fileext = ".csv")
feb <- tempfile(fileext = ".csv")
write.csv(mtcars[1:16, ], jan, row.names = FALSE)
write.csv(mtcars[17:32, ], feb, row.names = FALSE)
both <- tl_read(c(jan, feb), .quiet = TRUE)
table(both$source_file)
# A .txt file is read as CSV unless told otherwise
txt <- tempfile(fileext = ".txt")
write.table(mtcars, txt, sep = "\t", row.names = FALSE)
tl_read(txt, format = "tsv", .quiet = TRUE)
unlink(c(csv, jan, feb, txt))
Read from Google BigQuery
Description
Executes a SQL query against Google BigQuery and returns the result as a
tidylearn_data object. Requires the bigrquery package and
valid Google Cloud authentication.
Usage
tl_read_bigquery(project, query, dataset = NULL, ...)
Arguments
project |
Google Cloud project ID, or a
|
query |
A SQL query string (Standard SQL). |
dataset |
Optional default dataset for unqualified table names,
in |
... |
Additional arguments passed to
|
Value
A tidylearn_data object containing the query results.
Examples
## Not run:
# Needs Google Cloud credentials
tl_read_bigquery(
project = "my-project",
query = "SELECT * FROM `my_dataset.my_table` LIMIT 1000"
)
# Unqualified table names resolve against `dataset`
tl_read_bigquery(
project = "my-project",
query = "SELECT * FROM my_table LIMIT 1000",
dataset = "my_dataset"
)
## End(Not run)
Read a CSV file
Description
Reads a CSV file into a tidylearn_data object. Uses readr when
available for faster parsing, with a base R fallback.
Usage
tl_read_csv(path, ...)
Arguments
path |
Path to a CSV file. |
... |
Additional arguments passed to |
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Examples
path <- tempfile(fileext = ".csv")
write.csv(mtcars, path, row.names = FALSE)
tl_read_csv(path)
unlink(path)
Read from a DBI database connection
Description
Executes a SQL query against an existing DBI connection and returns
the result as a tidylearn_data object. The connection is not closed
by this function — the caller is responsible for managing the connection
lifecycle.
Usage
tl_read_db(conn, query, ...)
Arguments
conn |
A DBI connection object (e.g., from
|
query |
A SQL query string. |
... |
Additional arguments passed to |
Value
A tidylearn_data object containing the query results.
Examples
# RSQLite imports DBI, so both are available here
conn <- DBI::dbConnect(RSQLite::SQLite(), ":memory:")
DBI::dbWriteTable(conn, "cars", mtcars)
tl_read_db(conn, "SELECT mpg, cyl, hp FROM cars WHERE cyl = 6")
DBI::dbDisconnect(conn)
Read all matching files from a directory
Description
Scans a directory for files matching a pattern or format, reads each one,
and row-binds them into a single tidylearn_data object with a
source_file column identifying the origin of each row.
Usage
tl_read_dir(
path,
pattern = NULL,
format = NULL,
recursive = FALSE,
.quiet = FALSE,
...
)
Arguments
path |
Path to a directory. |
pattern |
Optional regex pattern to filter file names (e.g.,
|
format |
File format to read. If |
recursive |
Logical. Should subdirectories be scanned? Default is
|
.quiet |
Suppress messages. Default is |
... |
Additional arguments passed to the format-specific reader. |
Value
A tidylearn_data object with an additional
source_file column giving each row's file as a path below
path, such as "2024/sales.csv".
Examples
dir <- tempfile("sales_")
dir.create(file.path(dir, "2024"), recursive = TRUE)
write.csv(mtcars[1:16, ], file.path(dir, "jan.csv"), row.names = FALSE)
write.csv(mtcars[17:32, ], file.path(dir, "2024", "feb.csv"),
row.names = FALSE)
# Only the top level unless asked to recurse
tl_read_dir(dir, format = "csv")
# Files in subfolders are labelled by their path below dir
all_months <- tl_read_dir(dir, recursive = TRUE, .quiet = TRUE)
table(all_months$source_file)
# Or select files by a regular expression
tl_read_dir(dir, pattern = "^jan", .quiet = TRUE)
unlink(dir, recursive = TRUE)
Read an Excel file
Description
Reads an Excel file (.xls, .xlsx, or .xlsm) into a
tidylearn_data object. Requires the readxl package.
Usage
tl_read_excel(path, sheet = 1, ...)
Arguments
path |
Path to an Excel file. |
sheet |
Sheet to read. Either a string (the name of a sheet) or an integer (the position of the sheet). Defaults to the first sheet. |
... |
Additional arguments passed to |
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Examples
# readxl ships a workbook with one data set per sheet
path <- readxl::readxl_example("datasets.xlsx")
tl_read_excel(path)
tl_read_excel(path, sheet = "mtcars")
Read from GitHub
Description
Downloads a raw file from a GitHub repository and reads it into a
tidylearn_data object. Accepts either a full GitHub URL or a
owner/repo shorthand with a file path.
Usage
tl_read_github(source, path = NULL, ref = "main", ..., trust_rds = FALSE)
Arguments
source |
A GitHub URL or |
path |
Path to the file within the repository (required when
|
ref |
Branch, tag, or commit SHA. Default is |
... |
Additional arguments passed to the format-specific reader. |
trust_rds |
Logical. Read an |
Value
A tidylearn_data object containing the downloaded data.
Reading R serialisation files
An .rds or .rdata file is rebuilt with readRDS() or
load(), which recreate whatever R objects the file describes.
Read them only from a repository you trust, on any version of R.
Examples
## Not run:
# Downloads over the network
tl_read_github("user/repo", path = "data/file.csv")
tl_read_github("https://github.com/user/repo/blob/main/data/file.csv")
## End(Not run)
Read a JSON file
Description
Reads a JSON file into a tidylearn_data object. Expects the JSON to
represent tabular data (array of objects or similar). A file with the
.ndjson extension is read as newline-delimited JSON, one record
per line, and so is a .json file that does not parse as a single
document but does as one record per line. Requires the jsonlite
package.
Usage
tl_read_json(path, flatten = TRUE, ...)
Arguments
path |
Path to a JSON file. |
flatten |
Logical. Automatically flatten nested data frames? Default is
|
... |
Additional arguments passed to |
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Examples
path <- tempfile(fileext = ".json")
jsonlite::write_json(mtcars, path)
tl_read_json(path)
# Newline-delimited JSON: one record per line
lines <- tempfile(fileext = ".ndjson")
jsonlite::stream_out(mtcars, file(lines), verbose = FALSE)
tl_read_json(lines)
unlink(c(path, lines))
Read from Kaggle
Description
Downloads a dataset file from Kaggle using the Kaggle CLI and reads it into
a tidylearn_data object. Requires the Kaggle CLI to be installed and
configured (pip install kaggle).
Usage
tl_read_kaggle(source, file = NULL, dest = NULL, type = "dataset", ...)
Arguments
source |
A Kaggle dataset slug (e.g., |
file |
The specific file to read from the dataset, as a path within
it; a file the CLI saved under its base name is found too. If
|
dest |
Directory to keep the download in. The default is a fresh
per-dataset directory under |
type |
Either |
... |
Additional arguments passed to the format-specific reader. |
Value
A tidylearn_data object containing the downloaded data.
Examples
## Not run:
# Needs the Kaggle CLI and Kaggle credentials
tl_read_kaggle("zillow/zecon", file = "Zip_time_series.csv")
tl_read_kaggle("titanic", file = "train.csv", type = "competition")
tl_read_kaggle("https://www.kaggle.com/competitions/titanic/data",
file = "train.csv")
## End(Not run)
Read from a MySQL/MariaDB database
Description
Connects to a MySQL or MariaDB database, executes a SQL query, and returns
the result as a tidylearn_data object. Accepts either a connection
string or individual connection parameters. Requires DBI and
RMariaDB.
Usage
tl_read_mysql(
dsn,
query,
dbname = NULL,
user = NULL,
password = NULL,
port = 3306,
...
)
Arguments
dsn |
A MySQL connection string (e.g.,
|
query |
A SQL query string. |
dbname |
Database name (if not in |
user |
Username (if not in |
password |
Password (if not in |
port |
Port number (if not in |
... |
Additional arguments passed to |
Value
A tidylearn_data object containing the query results.
Examples
## Not run:
# Needs a running MySQL or MariaDB server
tl_read_mysql(
dsn = "localhost",
query = "SELECT * FROM my_table",
dbname = "mydb",
user = "myuser",
password = Sys.getenv("MYSQL_PWD")
)
## End(Not run)
Read a Parquet file
Description
Reads a Parquet file into a tidylearn_data object. Requires the
nanoparquet package.
Usage
tl_read_parquet(path, ...)
Arguments
path |
Path to a Parquet file. |
... |
Additional arguments passed to |
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Examples
path <- tempfile(fileext = ".parquet")
nanoparquet::write_parquet(mtcars, path)
tl_read_parquet(path)
unlink(path)
Read from a PostgreSQL database
Description
Connects to a PostgreSQL database, executes a SQL query, and returns the
result as a tidylearn_data object. Accepts either a connection string
or individual connection parameters. Requires DBI and RPostgres.
Usage
tl_read_postgres(
dsn,
query,
dbname = NULL,
user = NULL,
password = NULL,
port = 5432,
...
)
Arguments
dsn |
A PostgreSQL connection string (e.g.,
|
query |
A SQL query string. |
dbname |
Database name (if not in |
user |
Username (if not in |
password |
Password (if not in |
port |
Port number. Default is 5432. |
... |
Additional arguments passed to |
Value
A tidylearn_data object containing the query results.
Examples
## Not run:
# Needs a running PostgreSQL server
tl_read_postgres(
dsn = "localhost",
query = "SELECT * FROM my_table",
dbname = "mydb",
user = "myuser",
password = Sys.getenv("PGPASSWORD")
)
# The same connection as a URL, with the password still kept out of it
tl_read_postgres(
"postgres://myuser@localhost:5432/mydb?sslmode=require",
query = "SELECT * FROM my_table",
password = Sys.getenv("PGPASSWORD")
)
## End(Not run)
Read an RData file
Description
Reads an RData (.rdata or .rda) file
into a tidylearn_data object. Since RData files
can contain multiple objects, use the name
argument to specify which object to extract.
If name is NULL and
the file contains exactly one data frame, it is returned automatically.
Usage
tl_read_rdata(path, name = NULL, ...)
Arguments
path |
Path to an RData file. |
name |
Optional name of the object to extract from the RData file. If
|
... |
Currently unused. |
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Reading files you did not create
load() rebuilds whatever R objects the file describes, so read
only files from a source you trust. On R before 4.4.0 a crafted file
can run code as it is read (CVE-2024-27322). The remote readers
tl_read_github() and tl_read_s3() refuse .rds
and .rdata files on those versions unless
trust_rds = TRUE.
Examples
path <- tempfile(fileext = ".rdata")
cars <- mtcars
flowers <- iris
save(cars, flowers, file = path)
# With more than one data frame in the file, name the one to read
tl_read_rdata(path, name = "flowers")
unlink(path)
Read an RDS file
Description
Reads an RDS file into a tidylearn_data object. Uses base R
readRDS() — no additional packages required.
Usage
tl_read_rds(path)
Arguments
path |
Path to an RDS file. |
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Reading files you did not create
readRDS() rebuilds whatever R objects the file describes, so
read only files from a source you trust. On R before 4.4.0 a crafted
file can run code as it is read (CVE-2024-27322). The remote readers
tl_read_github() and tl_read_s3() refuse .rds
and .rdata files on those versions unless
trust_rds = TRUE.
Examples
path <- tempfile(fileext = ".rds")
saveRDS(mtcars, path)
tl_read_rds(path)
unlink(path)
Read from Amazon S3
Description
Downloads a file from an S3 bucket and reads it into a tidylearn_data
object. The file format is auto-detected from the key's extension, or can be
specified explicitly. Requires the paws.storage package and valid AWS
credentials.
Usage
tl_read_s3(source, format = NULL, region = NULL, ..., trust_rds = FALSE)
Arguments
source |
An S3 URI (e.g., |
format |
Optional format override for the downloaded file. If
|
region |
AWS region. If |
... |
Additional arguments passed to the format-specific reader. |
trust_rds |
Logical. Read an |
Value
A tidylearn_data object containing the downloaded data.
Reading R serialisation files
An .rds or .rdata object is rebuilt with readRDS()
or load(), which recreate whatever R objects the file describes.
Read them only from a bucket you trust, on any version of R.
Examples
## Not run:
# Needs AWS credentials
tl_read_s3("s3://my-bucket/data/sales.csv")
tl_read_s3("s3://my-bucket/data/results.parquet", region = "eu-west-1")
## End(Not run)
Read from a SQLite database
Description
Opens a SQLite database file, executes a SQL query, and returns the result
as a tidylearn_data object. The connection is automatically closed
when done. Requires DBI and RSQLite.
Usage
tl_read_sqlite(path, query, ...)
Arguments
path |
Path to a SQLite database file ( |
query |
A SQL query string. |
... |
Additional arguments passed to |
Value
A tidylearn_data object containing the query results.
Examples
# RSQLite imports DBI, so both are available here
path <- tempfile(fileext = ".sqlite")
conn <- DBI::dbConnect(RSQLite::SQLite(), path)
DBI::dbWriteTable(conn, "cars", mtcars)
DBI::dbDisconnect(conn)
tl_read_sqlite(path, "SELECT mpg, cyl, hp FROM cars WHERE cyl = 6")
unlink(path)
Read a TSV file
Description
Reads a tab-separated file into a
tidylearn_data object. Uses readr when
available for faster parsing, with a base R fallback.
Usage
tl_read_tsv(path, ...)
Arguments
path |
Path to a TSV file. |
... |
Additional arguments passed to |
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp.
Examples
path <- tempfile(fileext = ".tsv")
write.table(mtcars, path, sep = "\t", row.names = FALSE)
tl_read_tsv(path)
unlink(path)
Read data from a zip archive
Description
Extracts a zip archive to a temporary directory and reads the contents.
If the archive contains a single data file, it is read directly. If
multiple data files are found, they are row-bound with a source_file
column. Use the file argument to select a specific file from
the archive.
Usage
tl_read_zip(path, file = NULL, format = NULL, .quiet = FALSE, ...)
Arguments
path |
Path to a zip file. |
file |
Optional name of a specific file within the archive to read:
its path within the archive ( |
format |
Optional format override for the file(s) inside the
archive. Without |
.quiet |
Suppress messages. Default is |
... |
Additional arguments passed to the format-specific reader. |
Details
An archive with a member whose name could reach outside the extraction
directory – an absolute path, a drive letter on Windows, or any
.. component – is refused before anything is extracted.
Value
A tidylearn_data object (a tibble subclass)
with attributes tl_source, tl_format, and
tl_timestamp. The archive is extracted to a temporary directory
that is cleaned up automatically. If multiple data files are found,
a source_file column gives each row's member as its path within
the archive.
Examples
# readr ships a zip archive holding one CSV
archive <- readr::readr_example("mtcars.csv.zip")
tl_read_zip(archive)
# Name a member to read just that one
tl_read_zip(archive, file = "mtcars.csv", .quiet = TRUE)
Feature Engineering via Dimensionality Reduction
Description
Use PCA, MDS, or other dimensionality reduction as a preprocessing step for supervised learning. This can improve model performance and interpretability.
Usage
tl_reduce_dimensions(
data,
response = NULL,
method = "pca",
n_components = NULL,
...
)
Arguments
data |
A data frame |
response |
Response variable name (will be preserved) |
method |
Dimensionality reduction method: "pca" or "mds" |
n_components |
Number of components to retain, at most the number
the method computes: one per numeric column for PCA, and |
... |
Additional arguments for the dimensionality reduction method |
Value
A list with components:
- data
The transformed data frame with reduced-dimension columns and the response variable (if provided).
- reduction_model
The fitted tidylearn dimensionality reduction model.
- original_data
The original input data frame.
- response
The response variable name, or
NULL.
Examples
# Reduce dimensions before classification
reduced <- tl_reduce_dimensions(
iris, response = "Species",
method = "pca", n_components = 3
)
model <- tl_model(reduced$data, Species ~ ., method = "tree")
Run a tidylearn pipeline
Description
Run a tidylearn pipeline
Usage
tl_run_pipeline(pipeline, verbose = TRUE)
Arguments
pipeline |
A tidylearn pipeline object |
verbose |
Logical; whether to print progress |
Value
The input tidylearn_pipeline object with its
$results component populated. Results include
$processed_data (the training data after preprocessing),
$preprocessing_stats (the medians, modes, centres and scales
learned from the training data, replayed by
tl_predict_pipeline), $model_results (a named
list of per-model fits and metrics), $best_model_name,
$best_model (the winning tidylearn_model), and
$metric_values.
The training data is every row when validation = "cv": each
model is scored across the folds and then refitted on all of them.
With validation = "split" it is the training split alone, since
the models kept are the ones fitted there and scored on the test rows.
Examples
pipe <- tl_pipeline(iris, Species ~ .,
models = list(tree = list(method = "tree")),
evaluation = list(metrics = "accuracy", validation = "cv",
cv_folds = 2, best_metric = "accuracy"))
pipe <- tl_run_pipeline(pipe, verbose = FALSE)
Save a pipeline to disk
Description
Save a pipeline to disk
Usage
tl_save_pipeline(pipeline, file)
Arguments
pipeline |
A tidylearn pipeline object |
file |
Path to save the pipeline |
Value
Called for its side effect of saving to disk; returns
NULL invisibly.
Examples
pipe <- tl_pipeline(iris, Species ~ .)
tl_save_pipeline(pipe, tempfile(fileext = ".rds"))
Semi-Supervised Learning via Clustering
Description
Train a supervised model with limited labels by first clustering the data and propagating labels within clusters.
Usage
tl_semisupervised(
data,
formula,
labeled_indices,
cluster_method = "kmeans",
supervised_method = "tree",
...,
cluster_args = list()
)
Arguments
data |
A data frame |
formula |
Model formula. The response must be categorical: a
factor, character or logical column, or an expression that computes a
factor or character vector from a column, such as |
labeled_indices |
Indices of labeled observations |
cluster_method |
Clustering method for label propagation:
|
supervised_method |
Supervised learning method for the final
model (default: |
... |
Additional arguments for the supervised model |
cluster_args |
A named list of arguments for the clustering step,
such as |
Details
Labels are propagated by majority vote within each cluster, so the response must be categorical. The rows are clustered on the predictors the formula names, into as many clusters as the labelled rows hold classes. A labelled row whose label is missing takes no part in the vote. Rows in a cluster where no labelled observation carries a label have no label to take, and labelled rows whose own label is missing have none either; both are left out of training, with a warning giving the counts. The pseudo-labels keep the response's level order, so the second level stays the positive class.
Value
A tidylearn model object with additional class
"tidylearn_semisupervised", trained on pseudo-labeled data. The
model includes a semisupervised_info element with
labeled_indices, cluster_model, label_mapping,
and n_unlabelled_dropped, the number of rows left out for
having no label to train on.
Examples
# Use only 10% of labels
labeled_idx <- sample(nrow(iris), size = 15)
model <- tl_semisupervised(iris, Species ~ ., labeled_indices = labeled_idx,
cluster_method = "kmeans",
supervised_method = "tree"
)
Split data into train and test sets
Description
Split data into train and test sets
Usage
tl_split(data, prop = 0.8, stratify = NULL, seed = NULL)
Arguments
data |
A data frame |
prop |
Proportion for training set (default: 0.8) |
stratify |
Column name for stratified splitting. Each stratum is
split at |
seed |
Random seed for reproducibility |
Value
A list with two elements:
$trainA data frame containing the training subset.
$testA data frame containing the test subset.
A split that leaves the test set empty, as a single row does, warns.
Examples
split_data <- tl_split(iris, prop = 0.7, stratify = "Species")
train <- split_data$train
test <- split_data$test
Perform stepwise selection on a linear model
Description
Perform stepwise selection on a linear model
Usage
tl_step_selection(
data,
formula,
direction = "backward",
criterion = "AIC",
trace = FALSE,
steps = 1000,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the initial model |
direction |
Direction of stepwise selection: "forward", "backward", or "both" |
criterion |
Criterion for selection: "AIC" or "BIC" |
trace |
Logical; whether to print progress |
steps |
Maximum number of steps to take |
... |
Additional arguments to pass to step() |
Details
Every candidate model is fitted on the same rows: those with no
missing value in any variable of formula. When that leaves out
rows a candidate could otherwise use, a message gives the count, and
the fit records the rows in its na.action, as lm()
records the rows it drops itself.
Forward and both-direction selection start from a model holding the
intercept, or none if formula removes it with - 1, and
any offset() terms in formula. Neither is ever dropped.
Value
A tidylearn_model object of class
tidylearn_linear wrapping the selected lm
model. Access the underlying model via $fit and the selected
formula via $spec$formula. $data is the data passed in.
Examples
model <- tl_step_selection(mtcars, mpg ~ ., direction = "backward")
summary(model)
Stratified Features via Clustering
Description
Create cluster-specific supervised models for heterogeneous data
Usage
tl_stratified_models(
data,
formula,
cluster_method = "kmeans",
k = 3,
supervised_method = "tree",
...,
cluster_args = list()
)
Arguments
data |
A data frame |
formula |
Model formula |
cluster_method |
Clustering method: |
k |
Number of clusters |
supervised_method |
Supervised learning method (default:
|
... |
Additional arguments for the supervised models |
cluster_args |
A named list of arguments for the clustering step,
such as |
Details
The rows are clustered on the predictors the formula names, and a model is fitted to each cluster. A cluster whose rows all hold one class has nothing for a classifier to separate, so it gets no model and its rows are predicted as that class.
Value
A list with class "tidylearn_stratified" containing:
- cluster_model
The fitted clustering model.
- clusters
The training rows' cluster assignments.
- supervised_models
Named list of tidylearn models, one per cluster that holds more than one class.
- single_class_clusters
Named character vector giving, for each cluster whose rows all hold one class, that class.
- formula
The model formula.
- data
The original training data.
Examples
models <- tl_stratified_models(mtcars, mpg ~ ., cluster_method = "kmeans",
k = 3, supervised_method = "linear")
Create formatted tables for tidylearn models
Description
Dispatches to the appropriate table function based on model type and requested table type. Requires the gt package.
Usage
tl_table(model, type = "auto", ...)
Arguments
model |
A tidylearn model object |
type |
Table type (default: "auto"). For supervised models: "metrics", "coefficients", "confusion", "importance". For unsupervised models: "variance", "loadings", "clusters". MDS models are not supported. |
... |
Additional arguments passed to the underlying table function |
Value
A gt table object.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_table(model)
tl_table(model, type = "coefficients")
Formatted cluster summary table
Description
Produces a styled gt table showing cluster sizes and mean feature values for the columns the clustering used. Supports kmeans, pam, clara, dbscan, and hclust models.
Usage
tl_table_clusters(model, k = 3, digits = 2, ...)
Arguments
model |
A tidylearn clustering model object |
k |
For hclust models, the number of clusters to cut (default: 3) |
digits |
Number of decimal places (default: 2) |
... |
Additional arguments (currently unused) |
Value
A gt table object.
Examples
model <- tl_model(iris[, 1:4], method = "kmeans", k = 3)
tl_table_clusters(model)
Formatted model coefficients table
Description
Produces a styled gt table of model coefficients. Supports linear,
polynomial, logistic, ridge, lasso, and elastic net models. The numbers
come from tl_coefficients, which returns them as a tibble
if you would rather format them yourself.
Usage
tl_table_coefficients(
model,
lambda = "1se",
digits = 4,
conf_int = FALSE,
level = 0.95,
exponentiate = FALSE,
...
)
Arguments
model |
A tidylearn model object |
lambda |
For regularised models: |
digits |
Number of decimal places (default: 4) |
conf_int |
Whether to add a confidence interval (default:
|
level |
Confidence level for the interval (default: 0.95) |
exponentiate |
Whether to report odds ratios rather than log odds
(default: |
... |
Additional arguments (currently unused) |
Value
A gt table object.
See Also
tl_coefficients for the underlying tibble.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_table_coefficients(model)
tl_table_coefficients(model, conf_int = TRUE)
Compare multiple models in a formatted table
Description
Evaluates multiple tidylearn models and presents the results side-by-side in a styled gt table.
Usage
tl_table_comparison(..., new_data = NULL, names = NULL, digits = 4)
Arguments
... |
tidylearn model objects to compare |
new_data |
Optional test data for evaluation. If NULL, the models
are scored on their training data, which they must share: models
fitted on different data are an error asking for |
names |
Optional character vector of model names |
digits |
Number of decimal places (default: 4) |
Value
A gt table object. Its source note counts the
rows scored, per model when the models scored different rows.
Examples
m1 <- tl_model(mtcars, mpg ~ ., method = "linear")
m2 <- tl_model(mtcars, mpg ~ ., method = "lasso")
tl_table_comparison(m1, m2, names = c("Linear", "Lasso"))
Formatted confusion matrix table
Description
Produces a styled gt confusion matrix with correct predictions highlighted. Only available for classification models.
Usage
tl_table_confusion(model, new_data = NULL, ...)
Arguments
model |
A tidylearn classification model |
new_data |
Optional test data. If NULL, uses training data. |
... |
Additional arguments (currently unused) |
Value
A gt table object, with a row and a column for
each class the model was trained on. Rows of a class the model never
saw are left out, with a warning.
Examples
model <- tl_model(iris, Species ~ ., method = "forest")
tl_table_confusion(model)
Formatted feature importance table
Description
Produces a styled gt table of feature importance with a colour gradient. Supports tree-based, regularised, and xgboost models.
Usage
tl_table_importance(model, top_n = 20, digits = 2, ...)
Arguments
model |
A tidylearn model object |
top_n |
Maximum number of features to display (default: 20) |
digits |
Number of decimal places (default: 2) |
... |
Additional arguments (currently unused) |
Value
A gt table object.
Examples
model <- tl_model(iris, Species ~ ., method = "forest")
tl_table_importance(model)
Formatted PCA loadings table
Description
Produces a styled gt table of variable loadings on each principal component, with a diverging colour scale to highlight strong loadings.
Usage
tl_table_loadings(model, n_components = NULL, digits = 3, ...)
Arguments
model |
A tidylearn PCA model object |
n_components |
Number of components to show (default: all) |
digits |
Number of decimal places (default: 3) |
... |
Additional arguments (currently unused) |
Value
A gt table object.
Examples
model <- tl_model(iris[, 1:4], method = "pca")
tl_table_loadings(model)
Formatted evaluation metrics table
Description
Produces a styled gt table of model evaluation metrics from
tl_evaluate.
Usage
tl_table_metrics(model, new_data = NULL, digits = 4, ...)
Arguments
model |
A tidylearn supervised model object |
new_data |
Optional test data. If NULL, uses training data. |
digits |
Number of decimal places (default: 4) |
... |
Additional arguments passed to |
Value
A gt table object. Its source note counts the
rows scored.
Examples
model <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
tl_table_metrics(model)
Formatted PCA variance explained table
Description
Produces a styled gt table of variance explained by each principal component, with a colour gradient on cumulative variance.
Usage
tl_table_variance(model, n_components = NULL, digits = 4, ...)
Arguments
model |
A tidylearn PCA model object |
n_components |
Maximum number of components to show (default: all) |
digits |
Number of decimal places (default: 4) |
... |
Additional arguments (currently unused) |
Value
A gt table object.
Examples
model <- tl_model(iris[, 1:4], method = "pca")
tl_table_variance(model)
Test for significant interactions between variables
Description
Test for significant interactions between variables
Usage
tl_test_interactions(
data,
formula,
var1 = NULL,
var2 = NULL,
all_pairs = FALSE,
categorical_only = FALSE,
numeric_only = FALSE,
mixed_only = FALSE,
alpha = 0.05
)
Arguments
data |
A data frame containing the data |
formula |
A formula specifying the base model without interactions,
or a string that parses as one. |
var1 |
First variable to test for interactions |
var2 |
Second variable to test for interactions (if NULL, tests var1 with all others) |
all_pairs |
Logical; whether to test all variable pairs |
categorical_only |
Logical; whether to only test categorical variables |
numeric_only |
Logical; whether to only test numeric variables |
mixed_only |
Logical; whether to only test numeric-categorical pairs |
alpha |
Significance level for interaction tests |
Value
A data frame with one row per tested interaction pair, containing
columns var1, var2, p_value, significant
(logical), delta_r2 (change in R-squared), and
f_statistic, sorted by p_value ascending.
Examples
results <- tl_test_interactions(mtcars, mpg ~ wt + hp + cyl,
var1 = "wt", var2 = "hp")
Perform statistical comparison of models using cross-validation
Description
Perform statistical comparison of models using cross-validation
Usage
tl_test_model_difference(
cv_results,
baseline_model = NULL,
test = "t.test",
metric = NULL
)
Arguments
cv_results |
Results from tl_compare_cv function |
baseline_model |
Name of the model to use as baseline for comparison |
test |
Type of statistical test: |
metric |
Name of the metric to compare |
Value
A data frame with columns metric, model,
baseline, mean_diff, p_value, and
p_adj (Holm-adjusted p-value) containing pairwise
statistical comparisons against the baseline model. Each comparison
pairs the folds both models have a value for, and mean_diff is
the mean of those paired differences. With fewer than two such folds
the p-values are NA, with a warning.
Examples
m1 <- tl_model(mtcars, mpg ~ wt, method = "linear")
m2 <- tl_model(mtcars, mpg ~ wt + hp, method = "linear")
cv <- tl_compare_cv(mtcars, list(simple = m1, full = m2), folds = 3)
tl_test_model_difference(cv, baseline_model = "simple", metric = "rmse")
Transfer Learning Workflow
Description
Use unsupervised pre-training before supervised learning: the predictors the formula names are projected onto their principal components, and the supervised model is fitted on the component scores.
Usage
tl_transfer_learning(
data,
formula,
pretrain_method = "pca",
supervised_method = "tree",
...
)
Arguments
data |
Training data |
formula |
Model formula, or a string that parses as one. The PCA is fitted on the numeric predictors it names. |
pretrain_method |
Pre-training method. Only |
supervised_method |
Supervised learning method (default:
|
... |
Additional arguments passed to
|
Value
A list with class "tidylearn_transfer" containing:
- pretrain_model
The fitted dimensionality reduction model.
- supervised_model
The fitted supervised tidylearn model.
- formula
The model formula.
- method
The supervised learning method used.
Examples
model <- tl_transfer_learning(iris, Species ~ .,
pretrain_method = "pca", supervised_method = "tree")
Tune a deep learning model
Description
Tune a deep learning model
Usage
tl_tune_deep(
data,
formula,
is_classification = NULL,
hidden_layers_options = list(c(32), c(64, 32), c(128, 64, 32)),
learning_rates = c(0.01, 0.001, 1e-04),
batch_sizes = c(16, 32, 64),
epochs = 30,
validation_split = 0.2,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a
classification problem. |
|
List of vectors defining hidden layer configurations to try | |
learning_rates |
Learning rates to try (default: c(0.01, 0.001, 0.0001)) |
batch_sizes |
Batch sizes to try (default: c(16, 32, 64)) |
epochs |
Number of training epochs (default: 30) |
validation_split |
Proportion of the rows held out to score each configuration on, drawn at random (default: 0.2). Every configuration is scored on the same rows. |
... |
Additional arguments passed to keras's fit() for every
configuration; |
Value
A list with elements model (the best configuration refitted
as a tidylearn_model, so predict() and the deep plots take
it; the keras model is at $model$fit$model),
best_hidden_layers (optimal layer configuration),
best_learning_rate, best_batch_size, and
tuning_results (a data frame of all hyperparameter combinations
and their validation losses).
Examples
## Not run:
if (requireNamespace("keras", quietly = TRUE)) {
result <- tl_tune_deep(iris, Species ~ .,
hidden_layers_options = list(c(10), c(10, 5)),
learning_rates = c(0.01, 0.001), batch_sizes = c(32),
epochs = 5)
predict(result$model, iris[1:5, ])
}
## End(Not run)
Tune hyperparameters for a model using grid search
Description
Tune hyperparameters for a model using grid search
Usage
tl_tune_grid(
data,
formula,
method,
param_grid,
folds = 5,
metric = NULL,
maximize = NULL,
verbose = TRUE,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
method |
The modeling method to tune, one of the supervised methods
|
param_grid |
A named list of candidate values, one element for each
|
folds |
Number of cross-validation folds, a whole number between 2
and |
metric |
Metric to optimize: one of the names
|
maximize |
Logical; whether to maximize (TRUE)
or minimize (FALSE) the metric. |
verbose |
Logical; whether to print progress |
... |
Additional arguments passed to |
Value
A tidylearn model object fitted with the best hyperparameters.
Tuning results are stored as an attribute "tuning_results",
a list containing param_grid, results, best_params,
best_metric, metric, and maximize.
results has one row per evaluated combination: mean_metric
(the mean over the folds that produced a score), n_folds_ok (how
many of the folds did), and a column per parameter. A parameter
with a vector-valued candidate, such as hidden_layers, is a list
column. A fold produces no score when its fit fails, when the metric
is undefined on it – "auc" on a fold holding one class, or
"precision" on one where nothing is predicted positive – or
when none of its rows can be scored, as when every predictor is
missing there, which is warned about.
Only combinations with n_folds_ok equal to folds are
eligible to be best, since a mean over the folds that happened to
succeed is not comparable with a mean over all of them. If no
combination was scored on every fold, the best of those scored on the
most folds is used, with a warning saying whether fits failed or the
metric was undefined. If no combination was scored on any fold, the
function stops.
For method = "forest", an mtry above the number of
predictors is capped at that number, with a warning, and duplicate
combinations that result are evaluated once. Predictors are counted as
the forest is fitted on them, so a column removed with - id is
not one, and a matrix-valued term such as poly(hp, 2) is one per
column.
Examples
model <- tl_tune_grid(iris, Species ~ ., method = "tree",
param_grid = list(cp = c(0.01, 0.1), minsplit = c(10, 20)),
folds = 2, verbose = FALSE)
Tune a neural network model
Description
Tune a neural network model
Usage
tl_tune_nn(
data,
formula,
is_classification = NULL,
sizes = c(1, 2, 5, 10),
decays = c(0, 0.001, 0.01, 0.1),
folds = 5,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a
classification problem. |
sizes |
Vector of hidden layer sizes to try |
decays |
Vector of weight decay parameters to try |
folds |
Number of cross-validation folds (default: 5) |
... |
Additional arguments to pass to nnet(). |
Value
A list with elements model (the best fitted nnet
model), best_size (optimal hidden-layer size), best_decay
(optimal weight decay), and tuning_results (a data frame of all
parameter combinations and their cross-validated errors: the
misclassification rate for classification, the mean squared error for
regression).
Examples
tuned <- tl_tune_nn(iris, Species ~ .,
is_classification = TRUE,
sizes = c(2, 5), decays = c(0, 0.01), folds = 3)
tuned$best_size
tuned$best_decay
tuned$tuning_results
# The grid this searched, drawn as a heatmap
tl_plot_nn_tuning(tuned)
Tune hyperparameters using random search
Description
Tune hyperparameters using random search
Usage
tl_tune_random(
data,
formula,
method,
param_space,
n_iter = 10,
folds = 5,
metric = NULL,
maximize = NULL,
verbose = TRUE,
seed = NULL,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
method |
The modeling method to tune, one of the supervised methods
|
param_space |
A named list of parameter spaces to sample from, one
element for each
|
n_iter |
Number of random parameter combinations to try, a whole number of at least 1 |
folds |
Number of cross-validation folds, a whole number between 2
and |
metric |
Metric to optimize, as for |
maximize |
Logical; whether to maximize (TRUE)
or minimize (FALSE) the metric. |
verbose |
Logical; whether to print progress |
seed |
Random seed for reproducibility |
... |
Additional arguments passed to |
Value
A tidylearn model object fitted with the best hyperparameters.
Tuning results are stored as an attribute "tuning_results",
a list containing param_space, results, best_params,
best_metric, metric, and maximize.
results has one row per iteration: iteration,
mean_metric, n_folds_ok, and a column per parameter, as
described for tl_tune_grid. The best parameters are chosen
by the same rules, and mtry is capped the same way; duplicate
draws are kept, so there are always n_iter rows.
Examples
model <- tl_tune_random(mtcars, mpg ~ ., method = "tree",
param_space = list(cp = c(0.01, 0.1), minsplit = c(10, 20)),
n_iter = 3, folds = 2, verbose = FALSE)
Tune XGBoost hyperparameters
Description
Tune XGBoost hyperparameters
Usage
tl_tune_xgboost(
data,
formula,
is_classification = NULL,
param_grid = NULL,
cv_folds = 5,
nrounds = 1000,
early_stopping_rounds = 10,
verbose = TRUE,
...
)
Arguments
data |
A data frame containing the training data |
formula |
A formula specifying the model |
is_classification |
Logical indicating if this is a
classification problem. |
param_grid |
Named list of parameter values to try. NULL (default)
tries every combination of |
cv_folds |
Number of cross-validation folds (default: 5) |
nrounds |
Upper bound on boosting rounds per parameter set (default: 1000). Early stopping normally halts well short of it, so this is a ceiling rather than a target; lower it to cap the search. |
early_stopping_rounds |
Early stopping rounds (default: 10) |
verbose |
Logical indicating whether to print progress (default: TRUE) |
... |
Arguments |
Value
A tidylearn_model object (the refit on full data using the
best hyperparameters, built by tl_model so that it records
them in $spec$args) with an attribute "tuning_results"
containing a list with elements param_grid, results
(per-combination CV output), best_params, best_iteration,
best_score, and minimize.
Examples
if (requireNamespace("xgboost", quietly = TRUE)) {
# The default grid is every combination of the large xgboost grid
# without nrounds -- this many:
default_grid <- tl_default_param_grid("xgboost", size = "large")
prod(lengths(default_grid[names(default_grid) != "nrounds"]))
# Name a smaller one to see it run, and cap nrounds so early stopping
# has less ground to cover. nthread = 2 keeps xgboost from taking
# every core it is offered.
tuned <- tl_tune_xgboost(iris, Species ~ .,
param_grid = list(max_depth = c(2, 4)),
cv_folds = 3, nrounds = 20, verbose = FALSE, nthread = 2)
results <- attr(tuned, "tuning_results")
results$best_params
results$best_iteration
# tuned is an ordinary model, refit on all rows at those settings
predict(tuned, iris[1:5, ])
}
Get tidylearn version information
Description
Get tidylearn version information
Usage
tl_version()
Value
A package_version object containing the version number
Examples
tl_version()
Generate SHAP values for XGBoost model interpretation
Description
Generate SHAP values for XGBoost model interpretation
Usage
tl_xgboost_shap(model, data = NULL, n_samples = 100, trees_idx = NULL)
Arguments
model |
A tidylearn XGBoost model object |
data |
Data for SHAP value calculation (default: NULL, uses training data). The response column is not needed. |
n_samples |
Number of samples to use (default: 100, NULL for all) |
trees_idx |
Boosting rounds to include, as a run of consecutive
rounds counted from 1 such as |
Value
A data frame with one column of SHAP values per feature (the
columns of the model's design matrix), a BIAS column and a
row_id column giving the row of the (sampled) data each row
explains. For a multiclass model there is one block of rows per class,
told apart by a class column. The columns of data whose
names are not already taken – the response, and any factor predictor,
whose SHAP columns are named after its levels – are appended for
reference.
Examples
if (requireNamespace("xgboost", quietly = TRUE)) {
model <- tl_model(mtcars, mpg ~ ., method = "xgboost", nthread = 2)
shap <- tl_xgboost_shap(model, n_samples = 20)
}
Visualize Association Rules
Description
Create visualizations of association rules
Usage
visualize_rules(rules_obj, method = "scatter", top_n = 50, ...)
Arguments
rules_obj |
A tidy_apriori object or an arules rules object. A table of rules is refused: the plots need the rules object. |
method |
Visualization method: "scatter" (default), drawn by
tidylearn, or a method of arulesViz's |
top_n |
Number of rules to visualize, those with the highest lift (default: 50) |
... |
Additional arguments passed to plot() for rules visualization |
Value
For method = "scatter", a ggplot
object. Other methods return what arulesViz's plot()
returns: a ggplot object for "graph", "grouped" and "matrix", and for
"paracoord", which draws with grid, a grid vpPath.
Examples
if (requireNamespace("arules", quietly = TRUE)) {
data("Groceries", package = "arules")
res <- tidy_apriori(Groceries, support = 0.001, confidence = 0.5)
visualize_rules(res, method = "scatter")
}