Molecule guides
Reading Peptide Sequences: Codes, Modification Shorthand and Mass Arithmetic
Ten minutes of notation will let you check any vial against its certificate, predict its mass and anticipate how it behaves in solution.
Sequences run left to right, beginning at the N-terminus and ending at the C-terminus, written in one-letter or three-letter amino-acid codes, with each chemical modification noted as a prefix or suffix attached to the terminus it belongs to. Picking up this notation takes roughly ten minutes and gives you the quickest possible cross-check between a vial and its certificate of analysis: the sequence defines identity, allows the molecular weight to be predicted, and accounts for most of a molecule's behaviour once dissolved. What follows covers both code systems, the direction convention, the modification shorthand that actually appears on research labels, and a worked calculation converting a sequence into a mass you can hold against a COA. Everything described concerns laboratory reference material.
Why the N-to-C direction is not optional
Residues link through a peptide bond created between one residue's carboxyl group and the next residue's amino group. Because that reaction has a direction, each chain ends differently: one end keeps a free amino group (the N-terminus), the other a free carboxyl group (the C-terminus). Convention everywhere puts the N-terminus first. Gly-Glu-Pro and Pro-Glu-Gly are distinct molecules of identical composition and identical mass, and no mass spectrometer will tell them apart, which is why MS identity confirmation normally accompanies a declared sequence instead of replacing one.
Numbering runs the same way. A label reading HGH Fragment 176-191 uses positions counted from the N-terminus of the 191-residue parent protein, so the product is the 16 residues sitting at those positions rather than a purpose-designed sequence.
The two coding systems
Because three-letter codes are legible and unambiguous, they prevail on product labels and certificates. One-letter codes are compact, so they prevail in databases, journals and synthesis order forms. Both notate the same twenty proteinogenic residues.
Leucine and isoleucine are structural isomers and therefore weigh the same. That one fact underlies a familiar annoyance in peptide analytics: routine mass spectrometry cannot separate L from I, so telling them apart depends on the synthesis record and chromatographic behaviour rather than on mass.
Working through an actual label
Our catalogue gives BPC-157 as Gly-Glu-Pro-Pro-Pro-Gly-Lys-Pro-Ala-Asp-Asp-Ala-Gly-Leu-Val, or GEPPPGKPADDAGLV in single letters. That is fifteen residues, which is what "pentadecapeptide" means. The three prolines in a row near the N-terminus impose a rigid bend on the chain, and the two neighbouring aspartates carry negative charge at neutral pH. Both facts are legible directly from the sequence, and both bear on how the peptide behaves in solution.
Shorthand for modifications
Few research peptides are bare, unmodified chains. The notation is consistent once learned.
- A leading Ac- indicates acetylation at the N-terminus: an acetyl group caps the free amino group, adding 42.04 Da and cancelling a positive charge. TB-500 appears as Ac-Leu-Lys-Lys-Thr-Glu-Thr-Gln, or Ac-LKKTETQ.
- A trailing -NH2 indicates C-terminal amidation: -NH2 takes the place of the terminal -OH, removing 0.98 Da and a negative charge. Numerous natural signalling peptides are amidated in the body and usually need the amide for receptor activity; oxytocin is written Cys-Tyr-Ile-Gln-Asn-Cys-Pro-Leu-Gly-NH2.
- A D- prefix identifies the D-enantiomer, such as the D-Trp in GHRP-6 (His-D-Trp-Ala-Trp-D-Phe-Lys-NH2). D-residues weigh the same as their L-forms but make poor substrates for mammalian proteases, so chemists place them deliberately to slow degradation.
- Non-standard residues are spelled out. Aib denotes α-aminoisobutyric acid, which stabilises helices and resists proteases, used in Ipamorelin (Aib-His-D-2-Nal-D-Phe-Lys-NH2) and at position 2 of the GLP-1 backbone in semaglutide and tirzepatide. 2-Nal denotes 2-naphthylalanine.
- Brackets and superscripts flag substitutions against a parent sequence; [D-Ala²], for instance, means D-alanine occupies position 2.
- Cyclisation appears as a connecting line, a "cyclo(...)" wrapper or a declared Cys-Cys disulfide. In oxytocin, Cys1 and Cys6 bridge to close a six-residue ring, and that bond is sensitive to reduction, which matters in storage.
From sequence to molecular weight
To get molecular weight, total the residue masses, add 18.02 Da for the water lost when the chain forms, then correct for modifications. Two worked examples using catalogued figures:
- KPV (Lys-Pro-Val): 128.17 + 97.12 + 99.13 = 324.42; adding 18.02 gives 342.44 Da. Our catalogue shows 342.43, the 0.01 gap being rounding in the residue table.
- Epitalon (Ala-Glu-Asp-Gly, AEDG): 71.08 + 129.12 + 115.09 + 57.05 = 372.34; adding 18.02 gives 390.36 Da against a catalogued 390.35.
- Now add modifications. The same tetrapeptide acetylated and amidated would come to 390.36 + 42.04 − 0.98 = 431.42 Da.
Expect agreement to within about 0.1%. A gap of 18 Da normally means the water term was dropped; 42 Da means an acetyl group was overlooked; 1 Da means an amide was missed. Once the molecular weight is trustworthy you can move between mass and moles when preparing buffers, arithmetic covered in molecular weight, moles and molarity and in the molarity calculator.
Catalogue sequences versus certificate data
What a catalogue publishes is the intended sequence. What a COA publishes is measurement: an observed monoisotopic or average mass from MS and a purity value from reversed-phase HPLC. Neither method reads residues one by one unless MS/MS fragmentation data accompany it, which routine commercial certificates rarely include. Read the declared sequence as the manufacturer's claim and the measured mass as the test of that claim.
Frequent misreadings
- Reading from C to N. A reversed sequence is a different molecule, a retro-peptide, of identical mass, so always check that the leftmost residue really is the N-terminus.
- Counting a salt form as part of the sequence. Trifluoroacetate and acetate counter-ions are not residues; they add to the mass of the vial contents but not to the peptide. This is the usual reason a weighed quantity of powder holds less peptide than the label suggests, and why TFA content belongs on the COA.
- Mistaking fragment numbering for fragment length. "176-191" describes 16 residues, not 191.
- Overlooking the amide. A sequence active only when amidated is a different reagent from its free-acid counterpart, despite a difference of just one dalton.
- Treating a blend as one sequence. A co-formulated vial such as the KLOW Blend holds four sequences and has no single molecular weight; see blends vs single vials.
When the sequence, the stated molecular weight and the observed MS mass all line up, there is reasonable evidence the vial holds what the label claims. When they do not, the sequence is nearly always where the explanation is found fastest.
Questions
Which terminus comes first in a written sequence?
The N-terminus, the end bearing the free amino group, is always placed on the left, with the C-terminus on the right. Catalogues, certificates and the published literature all follow this convention. Writing the residues the other way round describes a different molecule of the same mass, so the direction is substantive rather than stylistic.
What does a trailing -NH2 signify?
C-terminal amidation, meaning the terminal carboxyl -OH has been swapped for an amide -NH2. That removes roughly 0.98 Da and one negative charge. Many secreted signalling peptides are amidated naturally, and published assays often show receptor activity disappearing without the amide, so amidated and free-acid forms count as separate reagents.
Why do D-amino acids appear in some sequences?
They weigh exactly the same as their L-counterparts but resist most mammalian proteases. Chemists position them at cleavage-prone sites to slow enzymatic breakdown in biological media. GHRP-6, written His-D-Trp-Ala-Trp-D-Phe-Lys-NH2, contains two.
Will a mass spectrometer give me the sequence?
Not from a mass measurement alone. A routine MS result shows that the total mass matches expectation, but it cannot separate leucine from isoleucine or a sequence from its reverse, since both share a mass. Residue-level information needs MS/MS fragmentation, which commercial certificates seldom provide.
How is molecular weight derived from a sequence?
Sum the average residue masses across the chain, add 18.02 Da for the water released during bond formation, then correct for modifications: +42.04 for an N-terminal acetyl, −0.98 for a C-terminal amide, −2.02 per disulfide bridge. Epitalon (Ala-Glu-Asp-Gly) gives 71.08 + 129.12 + 115.09 + 57.05 = 372.34, plus 18.02 = 390.36 Da.
Why does my figure differ a little from the catalogue value?
Discrepancies of a few hundredths of a dalton arise from rounding in residue tables or from mixing average with monoisotopic masses. Gaps of exactly 18, 42 or 1 Da indicate a missing water term, an unaccounted acetyl group or an overlooked amide. Bigger differences usually mean one figure includes the counter-ion salt and the other does not.
Does a blend have one sequence?
It does not. A co-formulated vial such as a KLOW or Wolverine preparation holds several different sequences and therefore several molecular weights. Its certificate should report every component separately, and molarity must be calculated per component from that component's stated mass in the vial.