What is CoreTrustSeal?
CoreTrustSeal is an international, community-based, non-governmental, and non-profit organization promoting sustainable and trustworthy data infrastructures. CoreTrustSeal offers a core-level certification based on the Core Trustworthy Data Repositories Requirements. This universal catalog of requirements reflects the core characteristics of trustworthy data repositories.
Immune Epitope Database and Analysis Resource has been certified as a Trustworthy Data Repository by the CoreTrustSeal Standards and Certification Board until 30 September 2026. The application is detailed below.
Description of Repository
The Immune Epitope Database and Analysis Resource (IEDB) is the most extensive, centralized resource for searching and analyzing data related to antibody and T cell epitopes for humans, non-human primates, rodents, and other animal species. Established in 2003 under the leadership of Dr. Alessandro Sette at the La Jolla Institute for Immunology, the IEDB (http://www.iedb.org) has been funded by the National Institute of Allergy and Infectious Diseases (NIAID), a component of the National Institutes of Health in the U.S. Department of Health and Human Services, since its inception. At its core, the IEDB is a freely available epitope database and prediction resource. An epitope, also referred to as an antigenic determinant, is the portion of an antigen that is recognized by the adaptive immune system. It is a chemical structure recognized by specific receptors - antibodies, major histocompatibility (MHC) molecules, and T cell receptors. As such, epitopes play critical roles in many diseases, including infectious diseases, autoimmune diseases, and allergies. They are also involved in organ transplants and blood transfusions. Epitopes can be characterized as continuous and discontinuous. Continuous or linear epitopes are sequences of amino acids. Discontinuous epitopes, sometimes called conformational, are composed of discontinuous segments of amino acids of one or more chains. The definitions of continuous and discontinuous also apply when one or more amino acids are replaced with a molecular entity that is not a peptide, such as a lipid. The study of epitopes has important applications in understanding the triggers of adaptive immune responses and in developing new vaccines, diagnostics, and therapeutics.
The IEDB has two avenues by which data is collected; biocuration of published, peer-reviewed articles and external data submissions from the research community. As of November 2020, the IEDB data has been derived from over 21,650 peer-reviewed journal articles. In addition, the IEDB contains over 300 direct submissions (corresponding to approximately 30% of the total data) from several NIH-funded large-scale epitope discovery programs and from researchers who directly approach the IEDB to deposit their data. This also includes negative data, which might typically appear in supplemental tables or might not be published at all. Slightly more than half of the references in the IEDB relate to infectious diseases, and about one quarter relate to autoimmune diseases, apart from HIV, which is captured separately in the Los Alamos HIV database. The remainder includes allergy, transplant, and other categories. Given that the data in the IEDB resides in the public domain, researchers can freely access, analyze, and publish work using this data. It should be stressed that the database exclusively contains information about epitopes derived from experiments. No predicted or model data resides in the database.
The curation of scientific literature started in 2004, requiring the curation of past and current relevant epitope literature in available peer-reviewed journals. As the IEDB has evolved, it has been necessary to change how biological concepts are captured in order to maximize accuracy. In addition, automated validation is continuously added. Consequently, there is a significant ongoing “recuration” effort of revising existing entries to improve data quality and consistency. The IEDB is now current with the published literature, and targeted PubMed queries are run biweekly, to make the data available in the IEDB within eight weeks of publication. Attached is the current organizational and team structure of the IEDB; a resource developed under the leadership of the La Jolla Institute for Immunology, consisting of 5 key teams with assigned leads, and 2 subcontractors for additional expertise.
Recent articles provide additional context for the IEDB more broadly:
- Vita R, Mahajan S, Overton JA, Dhanda SK, Martini S, Cantrell JR, Wheeler DK, Sette A, Peters B. The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Res. 2018 Oct 24. doi: 10.1093/nar/gky1006. PMID: 30357391; PMCID: PMC6324067.
- Fleri W, Vaughan K, Salimi N, Vita R, Peters B, Sette A. The Immune Epitope Database: How Data Are Entered and Retrieved. J Immunol Res. 2017;2017:5974574. doi: 10.1155/2017/5974574. Epub 2017 May 29. Review. PMID: 28634590; PMCID: PMC5467323.
Brief Description of the Repository’s Designated Community
The Immune Epitope Database and Analysis Resource (IEDB) is the single largest repository of immune epitope data in the world. It has an international user community, including biomedical researchers in academia, non-profit research institutes, and industry. The users include immunologists, microbiologists, virologists, and bioinformaticians who are interested in developing new vaccines, diagnostics, and therapeutics and who want to study the adaptive immune system. All data is freely available to users via http://www.iedb.org, and the data can also be downloaded in various user-friendly formats.
Level of Curation Performed
D. Data-level curation – includes conversion of data to new formats and enhanced documentation, and with additional editing of deposited data for accuracy
Comments
The IEDB data procurement requires systematic identification, categorization, curation, and quality‐checking processes, all of which have been documented in the publicly available IEDB curation manual. The curation manual is regularly updated by our Lead Ontology and Quality Manager to reflect the latest curation procedures.
The information curated in the IEDB is derived from the scientific literature cataloged in PubMed, and from direct data submissions from users, or more often from various NIH‐sponsored research efforts. Accordingly, for each epitope, it is necessary to extract the scientific details defining the assays in which the epitopes are defined and studied. The same epitope might be studied in multiple publications or submissions, and the same epitope is commonly tested in multiple assays, sometimes with different outcomes (e.g. a virus‐specific antibody might bind in an enzyme‐linked immunosorbent assay format but will not neutralize the live virus). Such nuances in the assay parameters accompanying the epitopes are objectively made available to view, based on the end‐user’s defined query.
A highly specialized expert curation process is pivotal to ensuring the consistency of such highly contextual data. This is achieved by PhD-level curators who review the data, and systematically extract and deposit the relevant information into a highly structured, computer‐operable format. For instance, the relevant information is not confined to a single location of the manuscript. Rather, the structure of the epitope as well as the contextual assay details may be reported in the methods section, while the data and interpretation of the results are often found in figures, tables, and text. The curator is, therefore, challenged to faithfully synthesize these disjointed elements of a publication into a concise format for user consumption. To ensure that the data is accurately represented, all curations are peer-reviewed by a second PhD-level curator to ensure accuracy and are updated as required. External data submissions are also reviewed and updated by curators to ensure the data is converted into the IEDB format. Of course, we require final approval from the external submitter prior to publishing the data.
Our curation process certainly adds value by enhancing the content to ensure it is understandable by users in a simple, tabulated form, rather than distributed through one or more manuscripts. Our curators also contact authors to obtain any additional information for data completeness (e.g. additional results, supplementary figures, etc.), which may not be within the user’s purview to do so.
The below recent article provides additional context to the curation process:
- Salimi N, Edwards L, Foos G, Greenbaum JA, Martini S, Reardon B, Shackelford D, Vita R, Zalman L, Peters B, Sette A. A behind-the-scenes tour of the IEDB curation process: an optimized process empirically integrating automation and human curation efforts. Immunology. 2020 Jul 2;161(2):139–47. doi: 10.1111/imm.13234. Epub ahead of print. PMID: 32615639; PMCID: PMC7496777.
Insource/Outsource Partners
The website and data reside at the La Jolla Institute for Immunology (LJI) in La Jolla, San Diego, California. There is a back-up site at the San Diego Supercomputer Center (SDSC), an organized research unit of the University of California, San Diego (UCSD). There is no SDSC certification, nor requirement to certify, that we are aware of. After further review, we have attached the SDSC SLA agreement. They have Colocation (COLO) service terms, which can be found on their website here: https://www.sdsc.edu/assets/docs/COLO_terms_of_service_2022.pdf. We have also attached the formal agreement between the La Jolla Institute and SDSC, which has been in place since 2010, and abides by these COLO terms.
Other Relevant Information
The IEDB websites typically receive a median of 20,000 visits per month, and this has grown dramatically in 2020 to 32,000 monthly visits based on Q1-3 data. This can likely be attributed to the growth in interest since the discovery of SARS-CoV-2 and use of the database and tools for research in this novel area. In 2020 thus far, the geographic breakdown of users (measured by visits to the main website) included Asia (39.6%), the Americas (36.7%), Europe (19.6%), Africa (2.2%), and Oceania (1.6%).
In terms of citations, the IEDB received 2,676 individual citations in 2019 (excluding self-citations), which is an increase of 462 citations from 2018. As of 2019, authors have been asked to cite one IEDB publication; “The Immune Epitope Database (IEDB): 2018 update”, which was published in Nucleic Acids Research. According to Google Scholar as of November 2020, this paper has been cited 267 times. Prior to this, authors were asked to cite one of three different papers published by the IEDB team in 2005, 2009, and 2014. According to Google Scholar as of November 2020, the original paper, “The Immune Epitope Database and Analysis Resource: From Vision to Blueprint”, has received 435 citations. The second paper, “The immune epitope database 2.0”, has received 655 citations. The third paper, “The immune epitope database (IEDB) 3.0”, has received 747 citations. In addition, many authors will cite the IEDB and/or its URL in their article without citing one of these three papers. We will complete the 2020 citation analysis in July 2021, and expect another considerable increase based on the growth of website visits in 2020. In addition, 72 US patent families cited or used the IEDB in 2019, which is 8 more than in 2018. The IEDB is projected to be cited by 84 US patent families in 2020, based on 56 records retrieved at the end of August 2020.
Overall, these statistics show that the IEDB has a global user base and its positive impact is far-reaching into the scientific community.
Organizational Infrastructure
R1 Mission/Scope
The goals of the IEDB are stipulated in the contract with the National Institute for Allergy and Immunology (NIAID). According to the Statement of Work, the contractor (LJI) is to:
- Maintain and further enhance a central web-based source of information on T cell epitopes and linear and conformational antibody/B cell epitopes (e.g., carbohydrates, lipids, and modified peptides) through curation of existing literature and direct submissions by the broader research community.
- Maintain and further enhance a central web-based source of data on ligand binding to MHC class I, class II, non-classical, and MHC-related molecules, including ligands shown experimentally not to bind to any of these molecules (i.e., negative binding data).
- Maintain and further enhance a central source of data on BCR and TCR repertoire information associated with T cell and antibody/B cell epitopes located within the IEDB.
- Foster further development of an Analysis Resource within the IEDB composed of more robust algorithms, mathematical models, and other predictive tools that support:
- Identification of novel antibody/B cell and T cell epitopes from genome or protein sequence information and predicting host responses to specific pathogens or immune-mediated diseases;
- Draw connections between both BCR and/or TCR repertoire sequence data, epitope binding and computational identification of epitopes from TCR/BCR sequence; and/or
- Facilitate identification of antibody and T cell epitopes associated with infectious or immune-mediated diseases for their use as targets for vaccine candidates and/or immune-based therapies.
Furthermore, the Statement of Work lists four additional components for fulfilling the aforementioned scope:
- Maintain, further develop, and improve the IEDB’s web-based relational database populated with antibody/B cell epitope and T cell epitope information. The IEDB will be freely accessible to the scientific community via an internet website and immune epitope information will be obtained primarily through curation of the scientific literature (relevant journal articles) and direct submissions from the broader research community.
- Maintain, enhance, further develop, and optimize the Analysis Resource for the IEDB. This includes online access to: (1) tools to help researchers locate and analyze information contained in the IEDB; (2) other relevant databases and related information; (3) data mining algorithms, mathematical models, and other sophisticated analytical tools to help researchers.
- Community Outreach activities to expand the user base and utility of the Immune Epitope Database resource for the broader research community.
- Interact with both current and future NIAID programs, which minimally include:
- Contractors supported by the B cell Epitope Discovery and Mechanisms of Action, Large-scale T cell Epitope Discovery, and the Allergen Epitope Discovery programs;
- Contractors supported by the Bioinformatics Integration Support Contract (BISC), Bioinformatics Resource Centers (BRCs) and the HIV Molecular Immunology Database.
It is a contractual obligation for the IEDB team to preserve and continue providing access to the data for the duration of our contract, and at the culmination of our contract, we will ensure all data is transferred to the incumbent (or back to NIAID) efficiently (see further details on this in R3).
Licenses
R2. The repository maintains all applicable licenses covering data access and use and monitors compliance.
Data contained in the IEDB website is within the public domain, free of all copyright restrictions, and made fully and freely available for both non-commercial and commercial use. IEDB data is manually curated, either from experimental data shown in publications or submitted datasets and the published data is linked to a PubMed identifier, and submitted data is linked to a submission identifier. All data is attributed to the publishing or submitting authors. This work is licensed under a Creative Commons Attribution 4.0 International License (Deed - Attribution 4.0 International - Creative Commons) and has been in place since May 2017. Users of the IEDB database are simply asked to cite the IEDB when using the resource, which can be found here - IEDB - Citation. More information on our Creative Commons license can also be found via the citation website, but, in summary, the license enables users to:
- Share — copy and redistribute the material in any medium or format
- Adapt — remix, transform, and build upon the material for any purpose, even commercially.
As stipulated in the Creative Commons license, these actions apply under the following terms:
- Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
- No additional restrictions — You may not apply legal terms or technological measures that legally restrict others from doing anything the license permits.
Further information on the IEDB’s use is available in the IEDB Terms of Use web page (IEDB - Terms of use).
The IEDB Analysis Resource (tools) is freely available to academic users through an open-source license, whilst commercial licenses are available to those wanting to utilize the tools within their private network. As agreed with the reviewers, this licensing scheme is outside the scope of this question, but more information can be found at this website if required - Download.
Continuity of Access
R3. The repository has a continuity plan to ensure ongoing access to and preservation of its holdings
The IEDB is in its third funding cycle with NIAID, which is a seven-year contract from December 2018 to December 2025. Prior to this, the IEDB was funded by a contract mechanism with NIAID for two additional funding cycles of eight and seven years. A fully developed transition plan has been a deliverable for all three contracts to enable a smooth transition of the database, and all respective user interfaces, from the incumbent to a new awardee or to the government, in the event of the incumbent not being selected for the contract renewal.
Hence, as stipulated in the Statement of Work in the subsection “Information Technology (IT) Resources, Facilities and Security”, the IEDB has a Continuity of Operations Plan, which includes a comprehensive Operational Recovery Plan (ORP) and Disaster Recovery Plan (DRP) that specifies the procedures used to restore operations following a natural or man-made disaster. This was submitted to NIAID in January 2019.
In agreement with the NIAID Contracting Officer, Emily Dubbaneh Bannister, and our IEDB Program Officer, Joseph Breen, we have now made the following documents publicly available in our Solutions Center:
- Operational Recovery Plan (ORP) and Disaster Recovery Plan (DRP) -
https://help.iedb.org/hc/en-us/articles/4406591519515--IEDB-Operational-Recovery-Plan-ORP-and-Disaster-Recovery-Plan-DRP- - IEDB Contract 3 Statement of Work - https://help.iedb.org/hc/en-us/articles/4406597263771-IEDB-Statement-of-Work-SOW-with-NIAID
Confidentiality/Ethics
R4. The repository ensures, to the extent possible, that data are created, curated, accessed, and used in compliance with disciplinary and ethical norms.
The data collected and distributed by the IEDB is considered public data and does not present ethical disclosure risks. The IEDB does not distribute author contact details in excess of what is already publicly available in PubMed; for example, curated literature contact information for the corresponding author is provided by the journal in which the article appears, hence is also available in the IEDB. Data submitted to the IEDB, as opposed to data curated from literature, are not promoted to the IEDB website for public access until the submitter releases them via a web interface. Therefore, submitters also review and approve contact details prior to publishing the data. This action must be performed by the submitter and cannot be performed by the IEDB team. The delay in the public release of data is typically done to allow time for the data to be published.
In regards to ensuring that deposited data was obtained under ethical conditions, the IEDB does not undertake any additional ethics clearance checks. This is because all externally submitted data is from NIH epitope contracts whose projects undergo ethical screening prior to data collection, especially when human or live specimens are in question. Therefore, there is no further assessment to be done by the IEDB. Similarly, when curating published literature from PubMed, these studies have already passed ethics approval and peer review, hence we do not perform additional checks.
Organizational Infrastructure
R5. The repository has adequate funding and sufficient numbers of qualified staff managed through a clear system of governance to effectively carry out the mission.
The repository is currently funded through a contract (75N93019C00001) from the National Institute of Allergy and Infectious Diseases (NIAID) to the La Jolla Institute for Allergy and Immunology (LJI). The IEDB is in its third funding cycle with NIAID, which is a seven year contract from December 2018 to December 2025. Prior to this, the IEDB was funded by a NIAID contract mechanism for two funding cycles of eight and seven years, therefore the repository has adequate funding and continuity.
NIAID is one of the 27 Institutes and Centers of the NIH, the largest funder of biomedical research in the world, and is widely recognized as a leader in the area of immunology research. NIAID research, in particular, strives to understand, treat, and ultimately prevent the myriad infectious, immunologic, and allergic diseases that threaten millions of human lives. LJI has extensive experience as a contractor organization working on NIAID grants and contracts. There is a clear governance structure from NIAID, with our Program Officer (PO), Dr. Joseph Breen, overseeing all major funding decisions regarding the IEDB. The IEDB leadership team also presents monthly updates to the PO and written report updates on IEDB goals on a quarterly basis to both the PO and Contracting Office (CO) at NIAID.
The IEDB’s scientific direction includes the leadership of Dr. Alessandro Sette, the Principal Investigator of the contract, who is a recognized leader in the area of immunology. He has considerable knowledge of immune epitopes, with a focus on the identification and biology of immune epitopes for infectious and immune-mediated diseases. In this respect, Dr. Sette has been the PI of several NIAID contracts for almost 30 years, including large-scale epitope identification contracts targeting smallpox/vaccinia virus, arenaviruses, dengue virus, mycobacterium tuberculosis, pertussis, and allergies. Dr. Sette has been the PI of the IEDB since its inception in 2003. Assisting him is Dr. Bjoern Peters, co-Principal Investigator, a bioinformatician who has been working on the IEDB since early 2004. His training in computer science, mathematics, and quantitative modeling, coupled with almost 20 years of working directly with clinicians, immunologists and biochemists, uniquely qualifies him to integrate the computational and experimental components of the IEDB. Therefore, the leadership team is well-qualified to lead the IEDB.
At the team level, there is a sufficient number of qualified staff managed through a clear governance structure to carry out the mission. The IEDB team includes PhD-level biocurators, bioinformaticians, database administrators, IT specialists, and project managers. The IEDB project is structured into 5 key teams; Curation (led by Dr. Alessandro Sette), Query & Reporting (led by Dr. Bjoern Peters), Tools (led by Dr. Bjoern Peters), IT Infrastructure (led by Dr. Jason Greenbaum - LJI Bioinformatics Core Director) and Outreach (led by the IEDB Project Manager). In addition, LJI has subcontracted the Technical University of Denmark (DTU) and Leidos Inc. to acquire complementary leadership, scientific and technical expertise.
Overall, the IEDB is both governed and comprised of highly skilled individuals, in a structured manner, to ensure that the mission is executed effectively. More information about the current IEDB team can be found here. The IEDB team is actively working to improve our support materials, including the acknowledgments. We are in the process of implementing a new support platform, Discourse, which will be available in 2024, and will increase the visibility of this information.
Expert Guidance
R6. The repository adopts mechanism(s) to secure ongoing expert guidance and feedback (either in- house, or external, including scientific guidance, if relevant).
The IEDB participates in an annual epitope meeting sponsored by NIAID that includes the NIAID large-scale epitope discovery contracts, which has ranged from 10-20 projects over the years. At this meeting, the IEDB presents a status update and future plans. Feedback from this collection of epitope experts is solicited to improve data quality, expand query and reporting features, and facilitate their data submission process. The meeting is also attended by NIAID staff with expertise in the field. In addition, since 2007 the IEDB team at LJI has published almost 30 meta-analyses of the data in the IEDB relating to a particular field, such as influenza A, tuberculosis, Ebola virus, and diabetes. These studies have created opportunities to interact with domain experts in specific fields of interest that have resulted in improved data quality and completeness.
The IEDB team also has access to the expertise of over 20 immunology faculty at LJI, immunological and disease experts at local research organizations,including Salk Institute, Scripps Research Institute, Sanford Burnham Prebys Medical Discovery Institute, and UC San Diego. Several team members are active participants on a variety of scientific advisory boards where they interact with domain experts in immunology, biocuration, ontology, and diseases, and bring back new ideas for enhancing the IEDB.
In addition, as part of our outreach efforts, IEDB staff attend scientific conferences and meetings to present information about the IEDB, its data, and its uses, and to gather feedback from the experts in attendance. The IEDB team also hosts an annual user workshop, which commenced in 2012. In addition to educating new and experienced users from a variety of backgrounds, one of the stated goals is to solicit feedback and comments on current and future features. These workshops have been a valuable source of new ideas for further development and prioritization for the team. The most recent 3 user workshops (2020-2022) can be seen in our Solutions Center via the following links:
- 2020 - https://help.iedb.org/hc/en-us/articles/360052475011-2020-IEDB-Virtual-User-Workshop-Presentations
- 2021 - https://help.iedb.org/hc/en-us/articles/4409650396571-2021-IEDB-Virtual-User-Workshop-Presentations
- 2022 - https://help.iedb.org/hc/en-us/articles/10097021647131-2022-IEDB-Virtual-User-Workshop-Presentations
All workshop recordings can be accessed via our IEDB YouTube channel. The latest information for the 2023 user workshops can be found here. Finally, the IEDB has also contracted usability engineering consultants to improve the user experience in accessing data.
Furthermore, in 2020, we established an IEDB Expert Committee, comprised of 17 power users of the IEDB database and tools. This group ranges from graduate and postdoctoral students to PI and NIAID-level representatives, specializing in both B and T cell research. We engage with this committee on a monthly basis, demonstrate new work, and solicit feedback to improve key features. Details of the Expert Committee can be found here in our Solutions Center (https://help.iedb.org/hc/en-us/articles/360057231112-IEDB-Expert-Committee), and it provides links to the 2 projects they have provided input on (IEDB Filter Options and the IEDB Query API - see links below).
- https://help.iedb.org/hc/en-us/articles/360053990892-What-are-the-IEDB-Filter-Options-
- https://help.iedb.org/hc/en-us/articles/4402872882189-Immune-Epitope-Database-Query-API-IQ-API
However, the IEDB team is actively working to improve our support materials, including the Expert Committee page. We are in the process of implementing a new support platform, Discourse, which will be available in the future, and will increase the visibility of this information. We will be updating the details of the members and including headshots, as well as a path to get involved with the group if desired. At this stage, the exact deliberations of the group are not published, but we can consider this in our updated support platform.
Lastly, as part of the IEDB contract, we host monthly teleconferences with our NIAID Program Officer, Dr. Joseph Breen. This provides an avenue to receive expert input and advice based on NIAID strategic priorities. This is imperative to ensure that we continue to meet NIAID expectations, as well as our users’ expectations.
Overall, we have multiple mechanisms to seek feedback, both internally and externally to the LJI team, and have avenues to access user feedback, strategic feedback from NIAID, and scientific input from other leaders in the field.
Links:
