Ambiguity residues B, X, Z and custom residues J, O, U
Mascot uses the IUPAC one-letter code for protein and peptide sequences. The 20 standard amino acids are encoded using the uppercase Latin letters ACDEFGHIKLMNPQRSTVWY. The remaining letters – B, J, O, U, X, Z – are also supported by Mascot and have a special meaning.
Ambiguity residues B, X, Z
The IUPAC one-letter symbols B, X and Z encode ambiguity: B is aspartic acid (D) or asparagine (N); X is “unknown”; and Z is glutamic acid (E) or glutamine (Q). Mascot interprets X as any of the 20 standard amino acids, which is mostly equivalent to “unknown”. These letters are used in public sequence databases like SwissProt, UniProt and NCBI nr, although they tend to be quite rare.
Amino acid frequencies can be found in your Mascot Server database status page. Scroll down to the database and click on Statistics. The text file lists a number of statistics, such as the total number of residues, total number of sequences and the AA frequency table:
Number of residues : 208906902 Number of sequences : 575503 Residue Frequency A 17247358 B 276 C 2902631 [...] X 8183 Y 6110160 Z 249
For UniProt, amino acid frequencies can also be found in the UniProtKB release statistics. In SwissProt 2026_02, above, the database has 39ppm of X, 1ppm of B and 1ppm of Z, which are found in about 0.01% of SwissProt sequences. In the whole of UniProtKB (including TrEMBL), the proportions of B and Z are similar but there is a lot more of X, 0.01%, found in about 0.05% of sequences.
During a database search, if Mascot encounters B, X or Z, it expands and tests all possibilities. For example, if the peptide sequence is PEBTIDE, Mascot tries to match PEDTIDE and PENTIDE to the MS/MS spectrum. To avoid excessive search times, it will test all possibilities up to a maximum of 3 B or Z residues and 1 X residue per peptide. The rules are summarised in Amino acid reference data.
The results file records the original sequence (PEBTIDE) and the residue substitution that yielded the peptide-spectrum match. Most of the time, when you’re viewing or exporting results, you’ll see the substituted or unambiguous sequence. The ambiguous sequence is shown in the Protein View report and also included in the mzIdentML export. Unfortunately, the other export formats don’t support reporting residue ambiguity.
Because B, X and Z are ambiguous and always substituted at match time, they cannot be used as a site for variable modification. The three remaining letters are different story.
Custom residues J, O, U
The letters J, O and U are special, because they are configurable. By default, U is selenocysteine, J is leucine (L) or isoleucine (I) and O is pyrrolysine. The composition is defined in Unimod. On your local Mascot Server, go to the Configuration Editor and Amino Acids, and it will show an Edit link next to J, O and U, as shown in the screenshot below:
The official IUPAC approval for O as pyrrolysine is as recent as 2009 (Biochemical Nomenclature Committee of IUPAC and NC-IUBMB newsletter 2009). In the 1980s, selenocysteine was encoded with X, although it quickly solidified to U, in 1999 (UPAC-IUBMB Joint Commission on Biochemical Nomenclature (JCBN) and Nomenclature Committee of IUBMB (NC-IUBMB) newsletter 1999).
The use of J for I/L is not standardised. In UniProtKB, the letter J is not used even once, while the database contains about 2ppm of U and <0.1ppm of O. Many years ago, the now-obsolete MSIPI database used J as an unconditional cleavage site, but few (if any) public sequence databases use it today. If you need to search peptides with a non-standard residue, co-opting J, O or U is a good way to do it.
J, O and U can be used freely as variable modification sites. Unimod includes a few standard modifications targeting selenocysteine, although nothing for pyrrolysine:
- Carbamidomethyl (U)
- Carboxymethyl (U)
- Oxidation (U)
- Dioxidation (U)
- MolybdopterinGD (U)
As always, you can add non-standard modifications in your local Mascot configuration with any chemistry.
Oxidised pyrrolysine (O+18) is a fairly common search parameter (for example, Hideki Yokoyama et al., 2026), while another recent paper (Chenfang Si et al., 2025) configured Mascot with carbamidomethyl of U and deselenation of U. And back in 2018, Bernd Thiede et al. demonstrated an intriguing way to encode glycopeptides using J, O and U for hexose, GlcNAc and sialic acid, respectively.
Non-standard residues in machine learning
Mascot Server 3.0 and later include MS2PIP for fragment intensity prediction, which can be used when refining the results with machine learning. MS2PIP doesn’t support any of the non-standard residues (B, J, O, U, X, Z). This doesn’t matter with B, X and Z, because Mascot always translates them into unambiguous residues before machine learning. However, any peptide match with J, O or U is filtered out when MS2PIP is active, and the correlation between predicted and observed spectrum is set to 0.
Because the proportion of peptides containing J, O or U is very small, this normally has no effect on the refined results. The other core features compensate for the lack of a predicted spectrum. But if your sequence database does contain a high proportion of selenocysteine or pyrrolysine, it’s best to not select any MS2PIP model.
Keywords: configuration files, modification, pyrrolysine, selenocysteine, Unimod