HLA-A*02:01 binding "AVGSYVYSV" at 1.90Å resolution
Data provenance
Information sections
- Publication
- Peptide details
- Peptide neighbours
- Binding cleft pockets
- Chain sequences
- Downloadable data
- Data license
- Footnotes
Complex type
HLA-A*02:01
AVGSYVYSV
Species
Locus / Allele group
Physicochemical Heuristics for Identifying High Fidelity, Near-Native Structural Models of Peptide/MHC Complexes.
There is long-standing interest in accurately modeling the structural features of peptides bound and presented by class I MHC proteins. This interest has grown with the advent of rapid genome sequencing and the prospect of personalized, peptide-based cancer vaccines, as well as the development of molecular and cellular therapeutics based on T cell receptor recognition of peptide-MHC. However, while the speed and accessibility of peptide-MHC modeling has improved substantially over the years, improvements in accuracy have been modest. Accuracy is crucial in peptide-MHC modeling, as T cell receptors are highly sensitive to peptide conformation and capturing fine details is therefore necessary for useful models. Studying nonameric peptides presented by the common class I MHC protein HLA-A*02:01, here we addressed a key question common to modern modeling efforts: from a set of models (or decoys) generated through conformational sampling, which is best? We found that the common strategy of decoy selection by lowest energy can lead to substantial errors in predicted structures. We therefore adopted a data-driven approach and trained functions capable of predicting near native decoys with exceptionally high accuracy. Although our implementation is limited to nonamer/HLA-A*02:01 complexes, our results serve as an important proof of concept from which improvements can be made and, given the significance of HLA-A*02:01 and its preference for nonameric peptides, should have immediate utility in select immunotherapeutic and other efforts for which structural information would be advantageous.
Structure deposition and release
Data provenance
Publication data retrieved from PDBe REST API8 and PMCe REST API9
Other structures from this publication
Data provenance
MHC:peptide complexes are visualised using PyMol. The peptide is superimposed on a consistent cutaway slice of the MHC binding cleft (displayed as a grey mesh) which best indicates the binding pockets for the P1/P5/PC positions (side view - pockets A, E, F) and for the P2/P3/PC-2 positions (top view - pockets B, C, D). In some cases peptides will use a different pocket for a specific peptide position (atypical anchoring). On some structures the peptide may appear to sterically clash with a pocket. This is an artefact of picking a standardised slice of the cleft and overlaying the peptide.
Peptide neighbours
P1
ALA
THR163
MET5
TRP167
TYR159
TYR59
GLU63
TYR171
TYR7
|
P2
VAL
PHE9
GLU63
VAL67
TYR7
HIS70
TYR99
MET45
LYS66
TYR159
|
P3
GLY
TYR159
TYR99
HIS70
LYS66
|
P4
SER
HIS70
ALA69
LYS66
|
P5
TYR
LEU156
THR73
HIS70
|
P6
VAL
HIS70
THR73
TYR99
ALA69
HIS114
ARG97
|
P7
TYR
TRP147
THR73
ASP77
GLN155
VAL152
LEU156
ARG97
|
P8
SER
TRP147
THR73
ASP77
VAL76
|
P9
VAL
TYR123
THR80
LEU81
TRP147
ASP77
TYR84
TYR116
LYS146
THR143
|
Colour key
Data provenance
Neighbours are calculated by finding residues with atoms within 5Å of each other using BioPython Neighboursearch module. The list of neighbours is then sorted and filtered to inlcude only neighbours where between the peptide and the MHC Class I alpha chain.
Colours selected to match the YRB scheme. [https://www.frontiersin.org/articles/10.3389/fmolb.2015.00056/full]
A Pocket
TYR159
THR163
TRP167
TYR171
MET5
TYR59
GLU63
LYS66
TYR7
|
B Pocket
ALA24
VAL34
MET45
GLU63
LYS66
VAL67
TYR7
HIS70
PHE9
TYR99
|
C Pocket
HIS70
THR73
HIS74
PHE9
ARG97
|
D Pocket
HIS114
GLN155
LEU156
TYR159
LEU160
TYR99
|
E Pocket
HIS114
TRP147
VAL152
LEU156
ARG97
|
F Pocket
TYR116
TYR123
THR143
LYS146
TRP147
ASP77
THR80
LEU81
TYR84
VAL95
|
Colour key
Data provenance
1. Beta 2 microglobulin
Beta 2 microglobulin
|
10 20 30 40 50 60
MIQRTPKIQVYSRHPAENGKSNFLNCYVSGFHPSDIEVDLLKNGERIEKVEHSDLSFSKD 70 80 90 WSFYLLYYTEFTPTEKDEYACRVNHVTLSQPKIVKWDRDM |
2. Class I alpha
HLA-A*02:01
IPD-IMGT/HLA
[ipd-imgt:HLA35266] |
10 20 30 40 50 60
GSHSMRYFFTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRMEPRAPWIEQEGPEYW 70 80 90 100 110 120 DGETRKVKAHSQTHRVDLGTLRGYYNQSEAGSHTVQRMYGCDVGSDWRFLRGYHQYAYDG 130 140 150 160 170 180 KDYIALKEDLRSWTAADMAAQTTKHKWEAAHVAEQLRAYLEGTCVEWLRRYLENGKETLQ 190 200 210 220 230 240 RTDAPKTHMTHHAVSDHEATLRCWALSFYPAEITLTWQRDGEDQTQDTELVETRPAGDGT 250 260 270 FQKWAAVVVPSGQEQRYTCHVQHEGLPKPLTLRWE |
3. Peptide
|
AVGSYVYSV
|
Data provenance
Sequences are retrieved via the Uniprot method of the RSCB REST API. Sequences are then compared to those derived from the PDB file and matched against sequences retrieved from the IPD-IMGT/HLA database for human sequences, or the IPD-MHC database for other species. Mouse sequences are matched against FASTA files from Uniprot. Sequences for the mature extracellular protein (signal petide and cytoplasmic tail removed) are compared to identical length sequences from the datasources mentioned before using either exact matching or Levenshtein distance based matching.
Downloadable data
Components
Data license
Footnotes
- Protein Data Bank Europe - Coordinate Server
- 1HHK - HLA-A*02:01 binding LLFGYPVYV at 2.5Å resolution - PDB entry for 1HHK
- Protein structure alignment by incremental combinatorial extension (CE) of the optimal path. - PyMol CEALIGN Method - Publication
- PyMol - PyMol.org/pymol
- Levenshtein distance - Wikipedia entry
- Protein Data Bank Europe REST API - Molecules endpoint
- 3Dmol.js: molecular visualization with WebGL - 3DMol.js - Publication
- Protein Data Bank Europe REST API - Publication endpoint
- PubMed Central Europe REST API - Articles endpoint
This work is licensed under a Creative Commons Attribution 4.0 International License.