The Monomers We Can’t Find: Building a Map of Non-Canonical Chemical Space
You Can Buy It. Can You Find It?
Written by Julie Owen, ProteinQure
This is semaglutide. The peptide drug that has transformed obesity treatment by copying a naturally occurring gut hormone. Its clinical success rests on a few edits to the natural sequence (shown in orange). Aib is a non-canonical amino acid (NCAA) that blocks cleavage by DPP-4 (dipeptidyl peptidase-4), an endogenous protease.

NCAAs are central to peptide therapeutics. They let us engineer metabolic stability, permeability and potency, and introduce novelty. Vendors supply thousands of Fmoc-protected, SPPS-ready building blocks. The absence of interpretable databases means most chemists will select from a familiar shortlist of NCAAs (D-amino acids, β-homologues, nor- and homo-variants and N-methylated monomers). We want to reason across the whole space chemically, not just by reputation, and answer the question every peptide chemist eventually arrives at:
“Which monomer should I test next?”
This led us to build the Peptide Monomer Database, a first-generation, free-to-use, searchable resource covering 2,488 monomers with SMILES in both protected and unprotected forms, InChIKeys, backbone classification across α/β/γ/δ/ε scaffolds, and natural-analogue mapping. Each entry carries vendor provenance across many commercial sources (including Bachem, Combi-Blocks, and WuXi TIDES), plus cross-references to PubChem, PDB, HELM and ChEMBL, where available. Each monomer has a chemically intuitive, ProteinQure-derived HELM-style shorthand and our popularity score, reflecting its appearance across external and vendor datasets, snapshotted July 2026.

The companion web application, Peptide Monomer Explorer, turns the database into an active discovery environment. A built-in 2D sketcher and fuzzy search allow you to locate monomers by drawn structure, SMILES or name, using partial or exact matching. For any entry, the Explorer returns Tanimoto-ranked nearest neighbours across the full collection.
The Explorer’s standout feature is the Chemical Exploration view. Every monomer is scored against one or more references on two independent axes: pharmacophore similarity on the x-axis and structural similarity on the y-axis. The useful monomers are the ones that match on one axis but not the other.
- Scaffold hops (high pharmacophore similarity, low structural similarity): bioisostere candidates with the same feature layout on a different scaffold.
- Shape-divergent (high structural similarity, low pharmacophore similarity): activity-cliff territory; a near-identical skeleton with a different feature profile.
- Close analogues: safe swaps.
- Distant monomers: least similar to the reference(s).

The thresholds are user-adjustable, which is deliberate. Where you draw the line between “close analogue” and “scaffold hop” is a scientific judgment, not a constant, and it should belong to the person asking the question rather than to us.
Structural similarity is the Tanimoto coefficient over Morgan fingerprints (radius 2); pharmacophore similarity is the Tanimoto coefficient over Gobbi 2D pharmacophore fingerprints. Physicochemical property plots sit alongside this for analytical work. Exports of the full database, or a filtered subset of it, are provided as CSV, TSV, SDF and JSON, with plots exportable as PNG.
HELM, via the Pistoia Alliance, gave the field a shared notation, but not a shared monomer set. The infrastructure supporting NCAAs hasn’t kept pace with their importance. Semaglutide is a blockbuster because someone knew to reach for Aib. The next Aib is almost certainly sitting in a vendor catalogue right now: Fmoc-protected, SPPS-ready, in stock. You can buy it. The only question is whether you can find it. That is what the Peptide Monomer Database and Peptide Monomer Explorer are for: turning “which monomer should I test next?” from institutional knowledge into a search.
Try the Peptide Monomer Explorer at https://monomers.proteinqure.com, no account needed.
The full database (v1.0.0) is published as a versioned, openly licensed dataset under CC BY-SA 4.0 on GitHub. Cite as:
ProteinQure. (2026). Monomer database datasets (Version v1.0.0) [Dataset]. Zenodo. https://doi.org/10.5281/zenodo.21684534Download it here to incorporate it into your own pipeline: https://zenodo.org/records/21684534
Found a monomer we’ve missed, or a record that’s wrong? Open an issue at https://github.com/ProteinQure/monomer-database-source or you can email us at monomers@proteinqure.com.