Reliability Limits and Decoding for Partial Nanopore Protein Rereads With Persistent State

2026-08-25Information Theory

Information TheoryEmerging Technologies
AI summary

The authors study how repeatedly reading the same protein through a nanopore might not give completely independent results and model this process as a special kind of communication channel with overlapping and synchronized reads. They show that certain inference methods get very close to the best possible predictions and analyze how errors scale when multiple reads are combined. Using simulated data, they validate their methods and identify limits on how much memory of past reads helps improve accuracy. Their results help understand the trade-offs between combining repeated measurements and the complexity of processing them.

nanopore sequencingfinite-alphabet channelBayes risknegative log-likelihoodimportance samplingposterior inferencesynchronizationcompound-pass datamemory effectssemi-synthetic data
Authors
Hongbin Ni, Haofan Dong, Ozgur B. Akan
Abstract
Repeated observations of one physical object need not constitute independent channel uses. We model partial nanopore protein rereads as a finite-alphabet channel with canonical content, persistent readout, and pass-local coverage and synchronization. For exact compound-pass data, matched inference approaches the equivalence-class canonical posterior, and sitewise excess Bayes risk admits an action-aware achievable exponent. In an aligned specialization, observation-local redraw can cause linear-in-$K$ growth in true-label negative log-likelihood (NLL). We derive order-$b$ projection-stability bounds and an exact passwise-fusion diagnostic. On a PASTOR-informed semi-synthetic hard-symbol channel, label-blind deterministic-mixture importance sampling (LB-IS) agrees with exact enumeration at $L=7$. At $L=24, K=10$, LB-IS meets every prespecified aggregate absolute marginal-posterior and score-agreement criterion against a fixed high-allocation reference in three selected conditions. Joint agreement holds for the representative and high-NLL conditions, while the near-zero condition remains inconclusive. Exact $L \leq 6$ benchmarks identify order 4 as the smallest tested common cap. At target scale, the reference supports selected unprojected functionals, while neither order 4 nor 5 attains joint agreement, defining a tested finite-memory boundary. Across 16 cells, the order-4 shared branch lowers NLL by 0.033-0.224 nats per residue relative to pass-local.