BioKERN: Biological Kernel Regularization for Histology-to-Transcriptomics Neighborhood Retrieval

2026-08-25Machine Learning

Machine Learning
AI summary

The authors developed BioKERN, a method that helps computers learn patterns in biological data by paying attention to how cells are arranged and related in space, not just exact one-to-one matches. BioKERN uses information about how close cells are and how similar their gene activity is to guide learning. When tested on mouse brain and human liver data, BioKERN better identified groups of similar cells compared to previous methods. The authors found that the improvement mainly comes from using this biological neighborhood information rather than making the model more complex. This suggests that including spatial biological context explicitly helps improve learning from biological data.

spatial biologytranscriptomicsembeddinginductive biasbiological kernelspatial proximitymultimodal learningVisiumembedding regularizationbiological neighborhood
Authors
Seungik Cho, Betul Orcan-Ekmekci
Abstract
Spatially resolved biology requires representations that preserve biological neighborhood structure rather than only exact cross-modal correspondences. Existing histology--transcriptomics objectives can emphasize instance-level matching even when non-paired spots share molecular or spatial context. We introduce BioKERN, a multimodal spatial representation-learning framework that incorporates biological structure as an explicit, learnable inductive bias. BioKERN constructs a training-time biological kernel by combining transcriptomic similarity and spatial proximity, then uses it to provide graded neighborhood supervision and regularize embedding geometry. Evaluation uses a fixed, model-independent biological neighborhood definition shared by all methods. Across Mouse Brain Visium and Human Liver GSE240429, BioKERN consistently improves biological-neighborhood retrieval over BLEEP in both single- and multi-scale settings. Controlled shared-architecture experiments show that most of the improvement arises from biological-kernel regularization rather than increased model capacity. These results support explicit biological geometry as an interpretable inductive bias for multimodal learning in spatial biology.