Noodle logo

Noodle

Biomedical literature discovery with source-backed entity mentions and provenance tracking for research questions.

AI-generated from publicly available materials.

Overview

Noodle is a biomedical literature discovery tool developed by Helena Bioinformatics, designed to help researchers, geneticists, and laboratory teams find and evaluate published scientific literature. It provides a public, read-only interface for submitting research questions and retrieving ordered sets of canonical publications, normalized identifiers, and source-backed entity mentions. Every result is linked to its source and provenance, keeping the boundary between retrieved evidence and clinical interpretation explicit. Noodle does not synthesize clinical conclusions, evaluate patient data, or replace professional review.

The tool draws on the Helena Literature Corpus, with PubMed supplying primary bibliographic records and Europe PMC and OpenAlex contributing additional publication coverage under defined source admission rules. Literature Mining, a component within the Helena infrastructure, owns corpus retrieval, deduplication, source convergence, ranking, and cursor continuity. Variant interpretation is handled separately by Folklore, a distinct component.

Search and Retrieval

  • Accepts natural language research questions as well as exact identifiers including PMID, DOI, PMCID, gene symbols, variant identifiers, phenotype terms, and OMIM identifiers.
  • Returns results in ranked order determined by semantic retrieval; the returned order is preserved as additional results load.
  • Semantic similarity explains why a record appears in search results but does not indicate evidence strength, causality, or clinical confidence.
  • Entities are linked only when a normalized identity or exact supported identifier is supplied by the source projection; unresolved text remains plain.
  • Gene and variant mentions are treated as source-indexed mentions unless a governed relationship supplies stronger evidence.

Publication Records and Provenance

  • Each record may carry a PMID, DOI, PMCID, canonical source URLs, author list, journal, and publication date.
  • Prepared publication pages also record a record version, manifest version, corpus timestamp, and source receipt.
  • Source-reported corrections and retractions remain visible within the interface; a materially changed prepared record receives a new version.
  • Freshness describes the source snapshot used to build a page and does not guarantee that the underlying science is current or correct.
  • A stale or unknown prepared record fails the indexing gate.

Integrated Genomic and Ontology Sources

  • Uses the Human Phenotype Ontology (HPO), release 2025-01-16, for phenotype term normalization, with HPO license and acknowledgement terms observed.
  • Incorporates project-generated Ensembl gene identifiers, symbols, and reference coordinates at release 113, presented with source attribution and release provenance; third-party assertions retain their own source labels.
  • Includes public ClinVar variant-classification data with submitter attribution preserved; ClinVar classifications are displayed as attributed source assertions and are never converted into Noodle clinical conclusions.
  • Includes primary data from gnomAD v4.1, presented with source attribution and version provenance; third-party annotations are kept under separate source labels.

Integration and Citation

  • Noodle exposes a Model Context Protocol (MCP) interface — the Noodle Biomedical Literature Discovery MCP — enabling use with AI agents; it provides public literature search, publication records, and bounded citation or semantic graph traversal through a single unauthenticated read-only endpoint.
  • The tool carries a persistent Research Resource Identifier: RRID:SCR_028920, suitable for citation in methods sections.
  • A version-specific DOI is available when a version-specific software citation is required alongside the RRID.

The public Noodle interface contains no patient, session, or operator data and does not evaluate patient context. It is intended as a traceable, source-grounded literature discovery resource where uncertainty remains explicit and conclusions remain open to revision as the underlying evidence changes.

Meta

Domain
Research Intelligence & Discovery
Subdomain
Scientific Literature Mining & Knowledge Discovery
Software type(s)
Database / Knowledge Base
Deployment type(s)
Cloud / SaaS
Industry vertical(s)
PharmaBiotechAcademic / ResearchDiagnostics / IVD
Development stage(s)
Research & DiscoveryPreclinical / Pre-MarketClinical
Target user(s)
Bioinformatician / Computational ScientistResearch ScientistClinical / Diagnostic Professional
Tag(s)
Uses AI