make_split_plan(group = , batch = )
directly?stratum column mean, and does splitGraph
stratify?make_split_plan(group = , batch = )
directly?For a dataset with one subject column and one batch column you
should: bioLeak’s make_split_plan() reaches the same
grouping in one call, and splitGraph adds a hop. splitGraph earns its
place when the structure is not one clean column per axis:
time_index that contradicts
timepoint_precedes: validate_graph() reports
these before any fold exists. Column-based grouping silently produces
folds from inconsistent metadata.mode = "composite" merges samples into connected components
over all chosen relations. bioLeak’s combined mode
instead picks a primary axis and deletes training rows sharing a
secondary level with the test set. Same metadata, different partitions;
splitGraph makes the choice explicit and reproducible.relatedness_edges_from_kinship() and
spatial_edges_from_coords() threshold them into edges and
the derivation takes connected components. No categorical column can
express this.split_spec records which relations, thresholds and strategy
produced each group, and travels as schema-checked JSON to Python or any
other consumer.Strict composite grouping is transitive closure. If sample A shares a subject with B, and B shares a batch with C, then A, B and C land in one group even though A and C share nothing directly. With a few large batches this collapses most of the dataset into one component:
meta <- data.frame(
sample_id = paste0("S", 1:6),
subject_id = c("P1", "P1", "P2", "P2", "P3", "P3"),
batch_id = c("B1", "B2", "B2", "B3", "B3", "B1"),
stringsAsFactors = FALSE
)
g <- graph_from_metadata(meta)
table(grouping_vector(derive_split_constraints(g, "composite", via = c("subject", "batch"))))
#>
#> component_1
#> 6Every subject bridges two batches, so all six samples form a single group and no split is possible. Three remedies, in order of preference:
spec$block_vars).strategy = "rule_based": each sample is grouped by
the first relation in priority that is available to it, so
relations do not chain.detect_dependency_components() to see the component
sizes before deriving, and summarize_leakage_risks() to see
which leakage paths a given mode actually severs (severed
column).grouping_vector(derive_split_constraints(g, "composite", strategy = "rule_based",
via = c("subject", "batch"),
priority = c("subject", "batch")))
#> S1 S2 S3
#> "composite_subject:P1" "composite_subject:P1" "composite_subject:P2"
#> S4 S5 S6
#> "composite_subject:P2" "composite_subject:P3" "composite_subject:P3"relatedness_edges_from_kinship(pairs, threshold) keeps a
pair when its kinship is at least the threshold;
spatial_edges_from_coords(coords, radius) keeps a pair when
its distance is at most the radius. The derivation then forms
connected components over the kept edges. Two consequences:
graph$metadata$edge_sources and
spec$metadata$threshold, so a reader of the spec can see
which cut produced the groups.pairs <- data.frame(id1 = c("P1", "P2"), id2 = c("P2", "P3"), kinship = c(0.26, 0.13))
meta <- data.frame(sample_id = c("S1", "S2", "S3"), subject_id = c("P1", "P2", "P3"))
build <- function(threshold) {
g <- build_dependency_graph(
list(create_nodes(meta, "Sample", "sample_id"), create_nodes(meta, "Subject", "subject_id")),
list(create_edges(meta, "sample_id", "subject_id", "Sample", "Subject", "sample_belongs_to_subject"),
relatedness_edges_from_kinship(pairs, threshold = threshold))
)
grouping_vector(derive_split_constraints(g, "relatedness"))
}
build(0.25) # only P1~P2 pass: {S1,S2}, {S3}
#> S1 S2 S3
#> "relatedness:component_1" "relatedness:component_1" "relatedness:component_2"
build(0.10) # P2~P3 also passes and chains: {S1,S2,S3}
#> S1 S2 S3
#> "relatedness:component_1" "relatedness:component_1" "relatedness:component_1"stratum column mean, and does splitGraph
stratify?No. stratum is an annotation: the outcome level
attached to each sample (from sample_has_outcome, or the
subject’s outcome via subject_has_outcome). It is exposed
through spec$stratum_var so a consumer such as
scikit-learn’s StratifiedGroupKFold or bioLeak’s
stratify = TRUE can balance folds. Balancing is execution
and belongs downstream; splitGraph only describes.
schema_version (currently 0.3.0) is independent of the
package version. It describes the on-disk JSON contract.NA and ignore unknown ones. A
differing major loads with a warning suggesting
migrate_dependency_graph_json() /
migrate_split_spec_json().inst/schema/<schema_version>/, and every written file
carries a $schema URL pointing at its own version, so the
reference stays valid after later bumps.
validate_graph_json() and
validate_split_spec_json() (or
read_*(validate = TRUE)) check a file against the installed
schema without a JSON Schema engine.The R reader is permissive at that boundary, which is worth seeing rather than taking on trust:
tiny <- data.frame(sample_id = c("S1", "S2"), subject_id = c("P1", "P2"),
stringsAsFactors = FALSE)
g_tiny <- graph_from_metadata(tiny)
p <- tempfile(fileext = ".json")
write_split_spec(as_split_spec(derive_split_constraints(g_tiny, "subject"),
graph = g_tiny), p)
# Pretend the file was written by a future splitGraph with a different major.
raw <- jsonlite::fromJSON(p, simplifyVector = FALSE)
raw$schema_version <- "1.0.0"
writeLines(jsonlite::toJSON(raw, auto_unbox = TRUE, null = "null"), p)
back <- withCallingHandlers(
read_split_spec(p),
warning = function(w) {
message("warning: ", conditionMessage(w))
invokeRestart("muffleWarning")
}
)
#> warning: Reading split_spec: JSON schema_version `1.0.0` differs in major version from installed splitGraph schema_version `0.3.0`. Loading anyway; consider `migrate_split_spec_json()` to upgrade the file.
class(back)
#> [1] "split_spec"
unlink(p)One asymmetry to know about if you consume the format outside R: the
shipped Python reader is stricter. Where R warns and loads,
splitspec raises a ValueError naming the
version and refuses the file, on the grounds that a non-interactive
consumer is better off failing than silently misreading a format it does
not know. Both implementations accept every schema sharing major
0.
Yes. graph_from_metadata() is an S3 generic, and the
SummarizedExperiment method reads colData() as
the metadata table. When colData has no
sample_id column the assay column names are used, so a
typical Bioconductor object needs no preparation at all. Pass
sample_id_col = to name a different column, and
columns = to map your own names onto the canonical ones
exactly as for a data frame.
meta <- data.frame(
sample_id = c("S1", "S2", "S3", "S4"),
subject_id = c("P1", "P1", "P2", "P2"),
batch_id = c("B1", "B2", "B1", "B2"),
stringsAsFactors = FALSE
)
se <- SummarizedExperiment::SummarizedExperiment(
assays = list(counts = matrix(0, nrow = 3, ncol = 4,
dimnames = list(NULL, meta$sample_id))),
colData = meta[, c("subject_id", "batch_id")]
)
g_se <- graph_from_metadata(se, graph_name = "from-se")
grouping_vector(derive_split_constraints(g_se, "subject"))
#> S1 S2 S3 S4
#> "subject:P1" "subject:P1" "subject:P2" "subject:P2"
# identical to building from the data frame directly
identical(
grouping_vector(derive_split_constraints(g_se, "subject")),
grouping_vector(derive_split_constraints(graph_from_metadata(meta), "subject"))
)
#> [1] TRUEEvery error inherits from splitgraph_error and carries a
code; see ?splitgraph_conditions for the
subclasses (splitgraph_schema_error,
splitgraph_reference_error,
splitgraph_ambiguity_error,
splitgraph_validation_error,
splitgraph_io_error).
Every step is linear in nodes plus edges. On a 20,000-sample
synthetic cohort the shipped benchmark
(inst/bench/pipeline.R) builds the graph, validates it,
derives a default composite constraint, enriches a spec and writes both
JSON files in about twelve seconds on a laptop. More than half of that
is writing the graph JSON (~6 s); the specification itself
writes in well under a second, and the derivation steps are hundredths
of a second. If you only need the handoff artifact, skip
write_dependency_graph().
Two caveats about the guard rail.
tests/testthat/test-performance.R is a wall-clock
budget at 5,000 samples with deliberately generous limits, not
a benchmark; only the composite derivation is additionally checked for
scaling, by timing 1,000 against 4,000 samples and failing if the ratio
exceeds 8 (a quadratic step would give roughly 16). Other steps could
therefore degrade somewhat without tripping it.
The one inherently quadratic output is the explicit sample-pair table
of detect_shared_dependencies() and
detect_dependency_components()$metadata$projection_edges,
whose size is the number of pairs sharing a target; grouping itself
never enumerates pairs.