
AINN-P1: Small Model, Big IntelligenceHow does a 167M-parameter protein foundation model outperform a 650M general-purpose model on the tasks that matter?
Ainnocence Inc.,an AI-driven drug discovery company, announced new results for AINN-P1, apurpose-built protein foundation model designed to deliver efficient andtransferable predictions for protein engineering and biologics discovery.
AINN-P1 takes a sequence-first approach,generating protein representations directly from amino acid sequences withoutrequiring multiple sequence alignments, structure prediction, or externalfunctional annotations. Conventional pipelines rely on homolog search,which is slow and requires large databases, while AlphaFold-class structureprediction still costs minutes per sequence. The encoder uses a multiplicativeLSTM rather than a Transformer. Where all-to-all attention requires n²connections and a key-value cache that grows with sequence length, recurrentstate passing requires n steps at a memory footprint that does not grow at alllinear O(n) time complexity against the Transformer's quadratic O(n²). The modelis trained autoregressively, predicting the next amino acid, rather than withthe masked language modeling objective used by the ESM family.
Transformer all-to-all attention(left) requires n² connections; mLSTM recurrent state passing (right) requiresn steps, with a memory footprint that does not grow with sequence length.
LeadingStability Prediction on ProteinGym
ProteinGym consolidates deep mutational scanning assaysacross four task categories: activity, binding, expression and stability.AINN-P1 recorded an average Spearman rho of 0.441 and a stability score of0.625, 6% above the structure-aware ProSST and 39% above the 100B-parameterxTrimoPGLM. Stability is a gating property for biologics developability, sincea protein that will not fold or survive manufacturing cannot proceed regardlessof its affinity.
Left: per-category performance, withAINN-P1 leading on stability. Right: accuracy vs. scale the 100B model sitsbelow the 167M model. AINN-P1 values are few-shot frozen-embedding results;baselines are published leaderboard values.
Across activity, expression and stability, the gap betweenAINN-P1 and the structure-aware ProSST is 0.03-0.06, and binding it is 0.02,indicating that sequence alone encodes more of the relevant constraint than iscommonly assumed.
TestingGeneralization on New Antibody Programs
The evaluation problem
Antibody discovery data are organized into programs. Eachprogram targets a specific antigen, and candidates within it come from the sameclonal families, sharing large stretches of framework sequence. Under a randomtrain/test split, close relatives land on both sides. A model can then scorewell simply by learning that a sequence belongs to a given program and that theprogram expresses well without capturing any of the biophysics that governsexpression.
Left: under a random split, everyprogram contains test points, so a model can win by recognizing programidentity. Right: under leave-program-out, an entire program is held out andonly transferable biophysics can help.
Ainnocence evaluated VHH single-domain antibody expressionunder both protocols, using proprietary data from real discovery programs.Under the random split the three encoders are nearly indistinguishable. Underleave-program-out the gap opens immediately.
The AUC drop is a direct leakagediagnostic. ESM2-base loses 0.222, meaning most of its random-split score camefrom program identity.
AINN-P1 exceeded the general-purpose 650M ESM2 by 0.15 AUCon new programs (0.810 vs. 0.660) with 3.9x fewer parameters and matched thefinetuned ESM2 within 0.006 AUC without any task-specific finetuning. To ruleout the downstream classifier as the driver, seven classifiers were evaluatedacross 21 encoder-by-classifier configurations; the ESM2-base column wasweakest at every classifier position, indicating the effect is a property ofthe embeddings rather than the head fitted on top of them.
“The most important result for us is not simply thatAINN-P1 performs well on a benchmark. It is that the model maintains strongperformance when evaluated on antibody programs it has not seen before. Thatdistinction is critical for drug discovery, where the real test is whether amodel can support decisions on the next program, not just reproduce patternsfrom the last one,” said Dr.Lurong Pan, Founder and CEO of Ainnocence.
About Ainnocence Inc.
Founded in 2021 and headquartered in California, Ainnocenceis a next-generation biotechnology company transforming drug discovery andsynthetic biology through AI-based, sequence-first engineering. The company’sself-evolving platform evaluates up to 10 billion molecules spanningproteins, antibodies, small molecules, nucleic acids, and chemical formulationswithin hours to weeks, enabling rapid, multi-objective design acrosstherapeutic, biological, and chemical systems. By reducing R&D timelinesand costs while increasing success rates, Ainnocence empowers industry andacademic partners to pursue complex biological innovation with greaterprecision and control.
Contact:
Dr.Lurong Pan, PhD
Founder and CEO
Ainnocence Inc.
+1 205-249-7424