Align a document-term matrix to a grid of fitted models
Source:R/utils.R
optop_align_dtm_to_models.RdEvery OpTop tool requires the document-term matrix and the models to
share one vocabulary in one order, because probabilities are aligned by
term name and never by position. When they drift apart (a feature was
trimmed after fitting, an engine reordered its vocabulary, models in
the grid were fit on slightly different preprocessing), the index
functions stop and point here. This helper subsets and reorders the
columns of dtm to the intersection of the models' vocabularies, in
the term order of the first model. Recompute the partition and the
baseline on the aligned matrix afterwards; they encode the support.
Arguments
- dtm
A document-term matrix of counts (rows are documents).
- models
A list of fitted topic models supported by
optop_as_theta_phi(), includingoptop_model()objects.
Details
Accepts a quanteda::dfm (returned as dgCMatrix) or any matrix with
column names.
See also
optop_make_partition() and optop_make_baseline(), which
must be recomputed on the aligned matrix.
Examples
rdirich <- function(n, k) {
g <- matrix(stats::rgamma(n * k, shape = 1), n, k)
g / rowSums(g)
}
theta <- rdirich(4, 2)
rownames(theta) <- sprintf("doc%d", 1:4)
phi <- rdirich(2, 6)
colnames(phi) <- sprintf("w%02d", 1:6)
m <- optop_model(theta, phi)
# counts whose columns are shuffled and carry one extra feature
counts <- matrix(rpois(4 * 7, 3), 4, 7,
dimnames = list(rownames(theta),
c("w03", "extra", "w01", "w06",
"w02", "w05", "w04")))
aligned <- optop_align_dtm_to_models(counts, list(m))
identical(colnames(aligned), m$terms)
#> [1] TRUE