Abstract
Major Histocompatibility Complex (MHC) molecules are central to vertebrate adaptive immunity, and MHC genes serve as key models in evolutionary genomics, offering insight into birth-and-death evolution, gene duplication, and the maintenance of genetic diversity. However, the organization and evolution of the MHC in species with giant genomes, such as salamanders, remain poorly understood. Here, we use comparative genomics, expression across multiple ontogenetic stages and tissues, as well as polymorphism data to investigate MHC evolution in newts. Contrary to earlier suggestions of a massively expanded MHC in salamanders, we find that the core MHC region remains relatively compact, demonstrating that genome gigantism does not scale proportionally in this region. Our finding also challenges the model of coevolution between a single classical MHC-Ia gene and antigen processing genes (APGs), revealing instead several polymorphic and highly expressed putative MHC-Ia located at varying distances from the APGs. MHC-I genes exhibit lineage-specific duplications and signs of concerted evolution, resulting in poorly resolved phylogenies. In contrast, MHC-II genes are more conserved and exhibit extensive trans-species polymorphism. Expression and polymorphism patterns identify putative nonclassical MHC-Ib genes, likely repeatedly derived from MHC-Ia genes, paralleling patterns seen in mammals but contrasting with the situation in fish and Xenopus frogs. In all seven studied species, some MHC-Ib genes show high relative expression during the larval stage but not at adulthood, suggesting a role in larval immunity. Our results underscore the importance of salamanders for understanding the evolution of complex regions in giant genomes and the architecture of the tetrapod MHC.
The Major Histocompatibility Complex (MHC) is a large genomic region originally identified as the genetic locus responsible for rapid allograft rejection in vertebrates—hence, its name (Klein 1986). Subsequent research revealed that these incompatibilities are primarily caused by a subset of genes now known as the classical MHC genes, which encode molecules essential to the molecular-level self/nonself recognition and the initiation of the adaptive immune response (Murphy and Weaver 2016). Because of its critical role in transplant rejection and its associations with susceptibility to both infectious and autoimmune diseases, the MHC has been a subject of intense work. Decades of research have found it to be among the most polymorphic and structurally variable regions in the genome, shaped by a complex dynamics of host-pathogen coevolution (for review, see Radwan et al. 2020). Beyond immunological significance, the MHC has thus become a prime model in evolutionary and comparative genomics, offering insights into birth-and-death processes, concerted evolution, gene conversion, coevolution of polymorphic genes, maintenance of genetic diversity, and immunogenetic trade-offs (Hughes and Nei 1988, 1992; Shiina et al. 2017; Kaufman 2018b; Migalska et al. 2019).
Classical MHC genes encode two types of cell surface glycoproteins that present peptide antigens to T cells: MHC class I (MHC-I, referred to as MHC-Ia) and MHC class II (MHC-II) (Murphy and Weaver 2016; Pishesha et al. 2022). Class I molecules are ubiquitously expressed and present cytosol-derived antigens to cytotoxic CD8+ T cells. In contrast, class II molecules are primarily expressed on antigen-presenting cells and present extracellularly derived peptides to helper CD4+ T cells. In addition to classical MHC, there are also nonclassical MHC genes (Adams and Luoma 2013). The nonclassical MHC-I (MHC-Ib) molecules are typically nonpolymorphic, have expression restricted to certain tissues, and often interact with innate immune cells or unconventional T cells (Mayassi et al. 2021). Some retain antigen-presenting capabilities, commonly for nonpeptide antigens, and often assume roles at the interface of adaptive and innate immunity, or even outside the immune system entirely (for a comprehensive review in mammals, see Adams and Luoma 2013). The presence and function of MHC-Ib in nonmammalian vertebrates are poorly studied, but large, divergent families of MHC-Ib genes have been identified in fish (Grimholt et al. 2015), and several Xenopus frog MHC-Ib genes have been characterized in detail, demonstrating roles in tadpole immunity (Edholm et al. 2013, 2018). Nonclassical MHC-II molecules, such as DM, generally function as chaperones that facilitate peptide loading onto classical MHC-II (Pishesha et al. 2022).
Whereas nearly all jawed vertebrates studied to date possess MHC, its genomic architecture is highly variable. Efforts to reconstruct the primordial organization of the MHC and trace the evolutionary trajectories leading to the configurations observed in extant vertebrates remain an area of active research (Ohta et al. 2006, 2019; Flajnik and Kasahara 2010; Kaufman 2018b; Veríssimo et al. 2023). One key insight from comparative studies is that the well-characterized genomic organization seen in mice and humans (the long-standing gold standard for MHC genomic research) is actually a derived state, specific to placental mammals. In humans, the MHC spans 4–5 Mb and contains well over 200 genes, many of which are not directly involved in immune response (Horton et al. 2004; Shiina et al. 2009). In humans—and more broadly in placental mammals—the Class I region, which includes MHC-Ia and some MHC-Ib genes, is separated from the Class II region by the Class III region. The Class III region contains both genes unrelated to immunity and immune-related genes with functions distinct from antigen presentation. The Class II region includes both classical and nonclassical MHC-II genes, as well as antigen processing genes (APGs), which are critical for preparing and transporting peptides for loading onto MHC-Ia molecules. At the edges of the MHC locus lie the “extended MHC regions,” among which the extended Class II region appears more conserved across species. In contrast, the ancestral MHC organization is now widely accepted to consist of a core MHC-I/II region (hereafter referred to as “core MHC region”), where all essential components of the “adaptive MHC” (Veríssimo et al. 2023) are located in close proximity: MHC-I genes and their APGs, as well as MHC-II genes. This arrangement occurs across a wide range of noneutherian vertebrates, including cartilaginous, sarcopterygian, and basal actinopterygian fish (Veríssimo et al. 2023), all three extant amphibian orders (He et al. 2023), lizards (Card et al. 2022), turtles (Bentley et al. 2023), crocodiles (He et al. 2022), birds (Kaufman et al. 1999), monotremes (Zhou et al. 2021), and even marsupials (Belov et al. 2006).
The architecture of the MHC region may have far-reaching evolutionary consequences. Specifically, the physical proximity—or separation—of genes influences the frequency of recombination between them, thereby potentially promoting or constraining coevolution among polymorphic genes that function together in a common pathway. This coevolutionary entanglement of interacting partners was proposed to limit their potential for gene duplication (Kaufman 1999; Ohta and Flajnik 2015; Martin and Kaufman 2022). A primary example is the coevolutionary theory concerning APGs and MHC-I. In placental mammals, a proposed translocation event placed a large block of genes, the class III region, between MHC-I and its associated APGs (which remained with class II), disrupting their physical linkage. This structural change is thought to have lifted the coevolutionary constraint, allowing for the emergence of a multigene MHC-Ia family and generalist APGs that provide multiple MHC-Ia molecules with broadly compatible, “average-best fit” peptides. However, the ancestral, tight linkage between MHC-Ia genes and their associated APGs, seen in noneutherian tetrapods, has been proposed to promote coevolution, while simultaneously constraining the expansion of MHC-Ia genes within the core MHC (for comprehensive reviews, see Kaufman 2015; Ohta and Flajnik 2015). Supporting this hypothesis, studies have found a single highly expressed MHC-Ia gene in the core MHC of both Xenopus and chicken, accompanied by more divergent clusters of MHC-Ib genes located further away—such as the XNC gene family in Xenopus and the Rfp-Y complex in chicken. This is at odds with growing evidence indicating that MHC-I in noneutherian taxa often undergoes expansion through gene duplication, likely including MHC-Ia genes. Whether this expansion generally occurs within the core MHC-I region remains unclear (Minias et al. 2019, 2022).
Salamanders (Urodela), one of the three extant amphibian orders, constitute a group that could be particularly informative in shedding light on many aspects of the MHC evolution, including those described above. As an early branching tetrapod lineage, salamanders are a natural choice for exploring the evolution of MHC architecture across the vertebrate phylogeny. Notably, they are known to frequently possess duplicated MHC-I genes (Minias et al. 2022), making them good candidates for tests of the APG-MHC-I coevolution hypothesis. However, the exceptionally large genome sizes, ranging from 10 to 120 Gb (https://www.genomesize.com), have long hampered genome-scale analyses in salamanders. Recent advances in sequencing technologies have yielded high-quality, chromosome-scale salamander genomes (Nowoshilow et al. 2018; Smith et al. 2019; Schloissnig et al. 2021; Brown et al. 2025). Although the complexity of the MHC region—and its rearrangements relative to the human—has introduced some confusion (for discussion, see Migalska et al. 2025), it is clear that we are now entering an era of unprecedented opportunity to resolve the MHC architecture in salamanders.
Apart from purely comparative aspects, such studies have a broad appeal, as several longstanding mysteries still surround salamander adaptive immunity. Similarly to other amphibians, they exhibit complex life cycles, often including an aquatic larval stage followed by metamorphosis (and typically a shift to a terrestrial lifestyle). Because of that, two major complications regarding the development of adaptive immunity emerge: (1) larvae face early pathogen exposure with limited time and cellular resources to develop full immune repertoires; and (2) metamorphosis may necessitate a repeated self-tolerization of new, adult-type tissues. Research in Xenopus proposed a reliance on MHC-Ib molecules during larval stages as a solution (Edholm et al. 2013, 2018). However, because the ontogenetic expression profiles of distinct MHC genes in Urodela remain largely unknown, it is unclear whether this mechanism could represent a universal feature tied to amphibian life history. A different molecular toolkit altogether could govern salamander immunity, as they have long been considered to exhibit subdued adaptive immune responses. Classical experiments reported slow (chronic) allograft rejection and weak mixed lymphocyte reactions in axolotls and several newt species, in contrast to anurans and other tetrapods (Cohen 1980; Pasquier et al. 1989; for review, see Kaufman and Volk 1994). Over the years, various theories have emerged to explain these subdued adaptive immune responses, often suggesting deficiencies in MHC function—at one point even referring to it as an “unrecognized MHC” (Kaufman et al. 1995). A detailed characterization of the structure and evolution of the salamander MHC would be a crucial first step toward resolving these controversies.
This work aims to present a comprehensive genomic and evolutionary view of the MHC in newts (Salamandridae, Pleurodelinae). We investigate the genomic organization of the MHC region, ontogenetic and tissue-specific expression profiles, and the evolutionary history of the MHC gene family. To achieve this, we integrated publicly available chromosome-scale genome assemblies from four newt species with newly generated long- and short-read transcriptome data from multiple developmental stages and tissues across seven species. Using comparative genomics, expression profiles, phylogenetic analysis, and structural modeling, we trace MHC gene duplications and signatures of adaptive evolution, and identify putative classical and nonclassical MHC genes.
Results
Comparative genomics of the MHC region(s) in newts
We analyzed four publicly available newt genomes (Lissotriton helveticus, L. vulgaris, Pleurodeles waltl, Triturus cristatus) (Fig. 1A), all assembled at the chromosome level. In each genome, there is a continuous MHC region assigned to a specific chromosome. Because the initial mapping of MHC transcripts to genomes indicated that automatic annotation of MHC genes was, in many cases, not accurate, MHC-I and MHC-II genes were annotated manually (Supplemental Data). The following description incorporates key observations informed by expression data and phylogenetic analyses, which are presented in detail later.
Newt phylogeny and a schematic representation of MHC genes and proteins. (A) Time-calibrated phylogeny of the studied species (Stewart and Wiens 2025); colored dots next to the species name indicate the availability of a reference genome. (B) Schematic representation of the structure of MHC molecules and their genes. MHC-I molecules are composed of a heavy chain with three extracellular domains (α1, α2, and α3), paired with a nonvariable β2-microglobulin (β2m) molecule, which is encoded by a gene outside the MHC region. MHC-II molecules are heterodimers, each consisting of an α and a β chain. Each chain contributes two extracellular domains: α1 and α2 for the α chain, and β1 and β2 for the β chain. The peptide-binding groove in MHC-I is formed by the α1 and α2 domains, whereas in MHC-II it is formed by the α1 and β1 domains. All MHC molecules also include a transmembrane domain and a cytosolic tail. There is a direct correspondence between structural domains and exons, as illustrated on the right-hand side. The first exon (gray box) encodes a leader peptide; each extracellular domain is encoded by a separate exon (indicated as “ex” with the exon number inside the box). The remaining regions—including the transmembrane domain and cytosolic tail—are encoded by a variable number of exons (black boxes), typically ranging from one to four, depending on the molecule type.

There are both considerable similarities and striking differences in the organization of the MHC region between the newt genomes (Figs. 2, 3). A core MHC region, hereafter defined as spanning from the first MHC-II to the last MHC-I in the vicinity of APGs or the last APG, is on an orthologous chromosome in all four species, although, because chromosome nomenclature is not standardized across assemblies, it is designated Chr 5 in L. helveticus and L. vulgaris, Chr 1 in T. cristatus, and Chr 6 in P. waltl (Fig. 2). In Lissotriton and Pleurodeles the MHC region is located at the terminal part of the chromosome reaching ∼ 0.5 (L. vulgaris)–9 (P. waltl) Mb from the end (Fig. 2). In T. cristatus, the MHC region is at the distance of >120 Mb from the end of the 2.5-Gb chromosome. Notably, the orientation of the region differs between assemblies, with the extended class II region located terminally/distally in L. vulgaris and T. cristatus and internally/proximally in L. helveticus and P. waltl. The availability of linkage maps of L. vulgaris and T. cristatus (France et al. 2025) allowed us to establish that the MHC region is actually located at the opposite ends of linkage group (LG) 1 in these two species, and a comparison with the P. waltl genome suggests that the region was translocated into a new location in the T. cristatus lineage. The length of the core MHC region varies from 5.5 Mb in P. waltl to 16.7 Mb in L. vulgaris.
The MHC region in amphibian genome assemblies and on linkage maps. The location and length of the MHC region comprising APG, MHC-II, extended MHC-II gene, and all annotated MHC-I genes located within 100 Mb. MHC genes in the four newt species were annotated manually in this study; manual annotation of MHC in A. mexicanum is from Migalska et al. (2025), and MHC genes in X. laevis are from the publicly available annotation of the v. 10.1 genome assembly. Note that, although in all newt genomes, the MHC region is on the same chromosome, naming and coordinates follow the original genome assemblies; numbers at the ends of chromosomes indicate coordinates (in Mb) in the original assemblies. Comparative linkage maps of L. vulgaris and T. cristatus (France et al. 2025) indicate that the MHC region is located on the opposite ends of LG1 in these species—cM 101 and 15, respectively.

Genes in the MHC region(s). Upper left: Genes in the genomic regions corresponding to those in Figure 2; APG, MHC-II, extended class II, and framework genes are indicated as colored blocks; MHC-I genes are shown individually; the width of the rectangles indicates gene lengths; filled rectangles below MHC-I genes show the maximum relative expression. Coordinates were rearranged so that the start/end of COL11A2, the extended class II gene adjacent in genomes to class II, is given coordinate 0. Pseudogenes are marked with asterisks. Upper right: MHC-I/MHC-like (if any) genes located outside the main MHC region. Lower part: The zoomed in “core” MHC region indicated with dashed rectangles in the upper plot. MHC-I, MHC-II and APG are marked and labeled individually, with the width of rectangles corresponding to gene lengths. Maximum relative expression: the maximum fraction of total MHC-I expression (FPKM) attributable to a sequence across all RNA-seq libraries from postmetamorphic individuals examined within a species (Supplemental Tables S3, S4).

In all species, multiple MHC-I genes (from three in L. vulgaris to eight in P. waltl), at least some of them highly expressed, are located between the MHC-II region and APGs (Fig. 3). In L. helveticus, most APGs are duplicated with four TAP2, three PSMB8, and two PSMB9 copies, whereas in three other species, each APG is a single copy gene. In both Lissotriton genomes, some highly expressed MHC-I genes are interspersed among APGs. Newt genomes contain also two divergent lineages of MHC-I sequences, which we name MHC-I-like1 and MHC-I-like2 (Fig. 4A; Supplemental Table S1). In both genomes where we found it (P. waltl and L. vulgaris), MHC-I-like1 gene is close to TAP1, and MHC-I-like1 pseudogene in the T. cristatus genome is also in this location. A single or two MHC-I-like2 genes are located between MHC-II and APGs in Lissotriton and Triturus but on a separate chromosome in Pleurodeles.
Phylogeny of newt MHC in the context of MHC sequences from other tetrapod taxa. The RAxML-NG maximum likelihood trees were constructed from protein sequences under the JTT+G4 amino-acid substitution model. The trees were rooted with zebrafish sequences (not shown). Support values for clades with the minimum bootstrap support of 70% (100 bootstrap replicates) are shown. The tree contains sequences from the G+T data set (as described in Methods) obtained from Pacific Biosciences (PacBio) Iso-Seq, de novo assembly of RNA-seq and protein sequences predicted for manually annotated MHC genes in genome assemblies of L. helveticus, L. vulgaris, P. waltl, and T. cristatus. Newt sequences are color-coded according to the species and species phylogeny is in Figure 1A. Additional salamanders included in the tree are the giant salamander (Andrias: GenBank AGY55962.1, AGY55988.1, AGY56015.1) and the axolotl (Ambystoma: sequences reported by Migalska et al. 2025). The trees include representatives of other major tetrapod groups, taken from GenBank or Ensembl (accessions starting with ENS): anurans (Xenopus: AAA16064.1, AAA16359.1, NP_001090513.1, BAA02842.1), birds (Gallus: ENSGALG00010003817, NP_001231990.1, ENSGALG00010003022), caecilians (Geotrypetes: XP_033779966.1, XP_033779817.1, XP_033779862.1), lepidosaurs (Sphenodon: ENSSPUG00000003482, ENSSPUG00000017679, ENSSPUG00000012680), and mammals (Homo: ENSG00000206503, ENSG00000204592, ENSG00000277263, ENSG00000196126). (A) MHC-I—a single (MHC-I-like1 and MHC-I-like2) or three (MHC-I) sequences randomly picked per newt species are included; for human and Xenopus, both MHC-Ia and MHC-Ib were included. (B) MHC-IIA, and (C) MHC-IIB—three sequences (if available) randomly picked per newt species were included.

In all species, some MHC-I genes are located outside of the core MHC region. Notably, MHC-I genes may be located on the same chromosome but >40 Mb (L. helveticus) or even >100 Mb (T. cristatus) from the core MHC region (Figs. 2, 3). In three species, there are some MHC-I genes at the opposite end of the MHC chromosome (L. vulgaris) or on other chromosomes (P. waltl and T. cristatus) (Fig. 3). In T. cristatus, there are four highly expressed MHC-I genes that occupy ∼3 Mb on Chr 5, not accompanied by any other genes usually found in the MHC region (Fig. 3). In the absence of another Triturus assembly or cytogenetic FISH evidence, we cannot completely rule out the possibility that the apparent translocation is an assembly error. However, genes flanking the additional MHC region from both sides map to a continuous genomic region in other newt assemblies, rendering misassembly unlikely.
In all four genomes, MHC-II genes are in the same order (Fig. 3): MHCIIB-IIA-BRD2-DMA-DMB and span from <1 Mb (P. waltl) to >3 Mb (L. vulgaris). In the L. vulgaris genome, there are two MHC-IIA and three MHC-IIB genes. The extended class II region is immediately adjacent to MHC-II in all species and spans from ∼ 5 Mb in Pleurodeles to >10 Mb in T. cristatus (Fig. 3).
Finally, we examined the immediate genomic context of the core MHC region, to see whether gene categories neighboring the MHC class I and II regions in mammals can be found in this group. Several framework genes (see Methods) were found in the vicinity of MHC in some species (Fig. 3). These are non-MHC genes interspersed in the eutherian MHC class I region (Amadou 1999), which are useful reference points for comparative analysis in mammals but not necessarily beyond (Belov et al. 2006). Additionally, some MHC class III genes were found within or close to the core MHC genes in T. cristatus, although we note that precise localization of Class III is beyond the scope of the present article.
Diversity of MHC sequences in newts
The MHC gene family frequently undergoes rapid expansions and contractions, leading to variation in gene content and blurred orthology among closely related species and even between haplotypes within species. As a result, a single reference genome per species is unlikely to capture the full extent of MHC diversity. Comprehensive sampling—both across individuals within species and across the phylogeny—is therefore key to understanding the evolution of MHC genes. To assess the diversity of MHC sequences in newts, we complemented the sequences annotated in genomes (hereafter referred to as data set G) with full-length coding sequences obtained from transcriptomes of, altogether, seven species (Supplemental Tables S2–S4; Fig. 1A). The coding sequences were obtained mostly through Pacific Biosciences (PacBio) Iso-Seq, supplemented by a selection of well-supported transcripts assembled from Illumina RNA-seq data. Translated sequences were clustered within each species to reduce redundancy (details are in Methods), and the longest protein from each cluster was selected as a cluster representative. This set of combined genome- and transcriptome-derived sequences will hereafter be referred to as data set G+T (Supplemental Data).
Phylogenies of MHC-I and MHC-II from the data set G+T are shown in Figure 4 and Supplemental Figures S1 and S2. The most striking feature of the MHC-I tree is the monophyly of all newt sequences and the presence among them of two well-supported clusters separated from the rest of the tree by long branches (Fig. 4A; Supplemental Fig. S1). These clusters, MHC-I-like1 and MHC-I-like2, are characterized by large deletions in the α1 and α2 domains, respectively (Fig. 5). MHC-I-like proteins, in particular MHC-I-like2, commonly lack conserved amino acids at the anchor residues,those important for anchoring the termini of antigenic peptides, that are conserved in the classical MHC-I of most taxa (Fig. 5; Supplemental Fig. S1; Supplemental Table S5). MHC-I-like1 sequences were detected in five species (I. apestris, L. boscai, L. vulgaris, P. waltl, and T. marmoratus), in each, most likely a single gene (Supplemental Table S6). We did not detect MHC-I-like1 sequences in either transcriptome data or in the genome assembly of L. helveticus, whereas in the T. cristatus genome, there is an apparent MHC-I-like1 pseudogene, which suggests the loss of functional MHC-I-like1 in these two species. MHC-I-like2 sequences, which, as long branches in the phylogeny indicate, evolve quickly, were detected in all species, and they do not cluster according to the species phylogeny (Fig. 4A; Supplemental Fig. S1; Supplemental Table S6). The number of variants per individual indicates multiple genes in some species. The remaining sequences on the MHC-I tree did not generally form well-supported clusters, although there was a tendency to group by species or genus (Fig. 4A; Supplemental Fig. S1). Given that MHC domains have different functions (Fig. 1B), they may be subject to different evolutionary pressures, including positive/purifying selection and gene conversion. Therefore, we additionally reconstructed separate phylogenies for MHC-I α1-3 domains (Supplemental Figs. S3–S5). These trees did not depart substantially from the full-length protein phylogenies, although branch lengths, also those leading to MHC-I-like1 and MHC-I-like2, were much shorter for α3 than for the other two domains. Additionally, in α3, the MHC-I-like2 clade was not strongly supported (Supplemental Fig. S5).
Protein alignments of α1–α3 domains, conserved and positively selected sites. Representative sequences of MHC-I, MHC-I-like1 (L1), and MHC-I-like2 (L2) proteins predicted from newt genomes were aligned to the Xenopus laevis allele with resolved crystal structure (Xela-UAAg, PDB: 6A2B) and representative, highly expressed MHC-I sequences of Ambystoma mexicanum (Amme-04 and Amme-05) (Migalska et al. 2025). Positively selected sites identified in newts by FUBAR are shaded in gray; conserved anchoring residues are marked with orange stars. Species abbreviations: (Lihe) L. helveticus, (Livu) L. vulgaris, (Plwa) P. waltl, (Trcr) T. cristatus.

The number of MHC-II sequences in the G+T data set was much smaller than that of MHC-I, and their phylogenies showed simpler patterns (Fig. 4B,C; Supplemental Fig. S2). In the case of nonclassical MHC-II, DMA phylogeny was poorly resolved, whereas the DMB tree generally reflected species phylogeny, with one or two sequences per species (Supplemental Fig. S2). Shorter DMA branch lengths indicate slower evolution compared to the DMB. Phylogenies of classical MHCIIA and MHCIIB showed considerable similarity to each other—there was a somewhat distinct Pleurodeles lineage and three divergent clusters, each of them containing sequences from all the remaining three genera. This pattern caused MHCIIA and MHCIIB phylogenies to depart drastically from species phylogeny. We did not find any evidence for the recently described MHC W type genes, that have both MHC-I and MHC-II features (Okamura et al. 2021), in newt genomes or transcriptomes.
The number of MHC genes and pseudogenes
All studied newts have multiple MHC-I genes (Supplemental Table S6). It is, however, not straightforward to estimate their number. Excluding MHC-I-like, there are from six (L. vulgaris) to 11 (T. cristatus) genes in the genome assemblies. The number of sequences retrieved from transcriptomes of individuals with Iso-Seq data available (these estimates include both Iso-Seq and Illumina RNA-seq G+T data set sequences in these individuals) varied from 10 in P. waltl to 15 in I. alpestris, and in all cases was larger than the number of genes detected in the assembly (Supplemental Table S6). This is not surprising as divergent alleles found in heterozygous individuals could have been assigned into separate clusters. Also, the number of MHC-I pseudogenes and gene fragments differs between species, ranging from eight in P. waltl to 23 in L. helveticus, with an overwhelming majority of pseudogenes and gene fragments within or in the vicinity of the MHC region (Fig. 3; Supplemental Table S6).
In each genome, there are single DMA and DMB genes and, except in L. vulgaris, single MHC-IIA and MHC-IIB genes (Fig. 3). We did not identify DM pseudogenes in any genome, and generally, there were none or few MHC-II pseudogenes, the only exception being 21 MHC-IIA pseudogenes in the T. cristatus genome. In contrast to MHC-I, MHC-II pseudogenes were located mostly outside the MHC region (Supplemental Table S7).
Ontogenetic and tissue expression profiles, polymorphism, and identification of putative nonclassical MHC-I genes
Ontogenetic profiles of MHC expression can shed light on the timing of adaptive immunity onset in newts. Moreover, when combined with expression pattern in multiple tissues and polymorphism data, they can help to distinguish classical MHC-Ia from nonclassical MHC-Ib genes. Overall, changes in expression of MHC-I and MHC-II were synchronized throughout development (Fig. 6). In P. waltl, there is a steady increase of MHC expression through the larval stages, whereas in the remaining six species, MHC expression in the larval period and immediately after metamorphosis is at least an order of magnitude lower than in adults (Fig. 6), with a sharp increase visible only in mature individuals. The ontogenetic expression profile of MHC-I-like genes is very different; they are expressed at much lower level, and their expression remains low throughout the larval period and metamorphosis and in adults (Fig. 6).
Overall MHC expression through ontogeny. Expression in fragments per kilobase of transcript length per million reads mapped (FPKM) of all MHC-I, MHC-II (MHC-IIA+MHCIIB), and MHC-I-like1 and MHC-I-like2 sequences through the ontogeny. Points are expression values (summed over all transcripts of a given MHC category) for RNA-seq libraries (Supplemental Table S4), and bars are averages across libraries for each species and stage. Species abbreviations: (Ia) I. alpestris, (Lb) L. boscai, (Lh) L. helveticus, (Lv) L. vulgaris, (Pw) P. waltl, (Tc) T. cristatus, (Tm) T. marmoratus.

In both G and G+T data sets, only some sequences are highly expressed at any larval stage or in adults. Whereas in Triturus and perhaps also P. waltl and L. vulgaris, not more than two MHC-I genes are highly expressed in adults, in I. alpestris, L. boscai, and L. helveticus, several genes appear to have similarly high expression (Fig. 7; Supplemental Fig. S6). Moreover, many sequences highly expressed in at least one tissue/ontogenetic stage had conserved amino acids in most (eight or nine out of nine) anchor residues—one of the hallmarks of MHC-Ia molecules (Supplemental Fig. S1; Supplemental Table S5). We did not detect pronounced, several-fold differences between adult tissues in relative expression of particular MHC-I sequences, with the exception of P. waltl Plwa-03 and Plwa-04 which have much higher relative expression in skin (tailtip) and T. marmoratus Tm-13 with much higher expression in lungs than in other tissues (Supplemental Fig. S7).
Expression through ontogeny of MHC-I genes annotated in the genomes (data set G). Relative expression, that is, the fraction of all RNA-seq reads in a library mapping to all MHC-I genes and pseudogenes in the assembly that mapped to a particular gene sequence, is shown. Box plots show medians, interquartile, and total ranges. Gray stripes indicate genes that show higher relative expression in larvae than in adults—those may be nonclassical (MHC-Ib) genes. Species abbreviations: (Lh) L. helveticus, (Lv) L. vulgaris, (Pw) P. waltl, (Tc) T. cristatus.

A striking pattern of ontogenetic changes in the relative expression of some MHC-I genes was detected in all seven species. During the larval period, when the overall MHC expression is low, the relative expression of some MHC-I genes was much higher than in adults (Fig. 7; Supplemental Fig. S6). Such a pattern suggests a possible functional divergence from the role played by classical MHC genes and hints at a role in larval immunity or development. The sequences with high relative expression in larvae of Triturus and Pleurodeles are easily identified visually in Figure 7 and Supplemental Figure S6, whereas the assignment of sequences to this category in other species is less straightforward. Still, such sequences do not form separate clades and occur in different parts of the phylogeny (Supplemental Fig. S8). Therefore, regardless of difficulties in defining the exact set of such sequences, it is clear that they have emerged repeatedly and are probably undergoing fast turnover on evolutionary time scales.
For the four species with assembled genomes, we estimated polymorphism of MHC-I genes by mapping variants obtained in previous amplicon-based studies (Palomar et al. 2021; Gaczorek et al. 2023) to the genome and calculating the number of variants mapping to a particular MHC-I gene (Supplemental Tables S8 and S9; Supplemental Fig. S9). Combining expression patterns with these polymorphism estimates (Fig. 8) suggests that sequences highly expressed in adults, which are also highly polymorphic, represent classical (MHC-Ia) genes. Sequences with high relative expression in larvae, lower expression in adults and/or substantial differences in expression among adult tissues, which also show low polymorphism (Fig. 8), likely represent nonclassical (MHC-Ib) genes. The status of the remaining genes, which have low relative expression in all examined stages and tissues, as well as generally low polymorphism, is unclear.
Polymorphism of MHC-I genes annotated in the genomes. Bar heights show polymorphism of each gene expressed relative to the most highly polymorphic gene in the species; MHC-I-like genes are not shown because reliable polymorphism estimates were not available. Polymorphism was estimated based on mapping amplicon variants detected in population samples to genes, as described in the text. Maximum relative expression: the maximum fraction of total MHC-I expression (FPKM) attributable to a sequence across all RNA-seq libraries from postmetamorphic individuals within a species. Species abbreviations: (Lh) L. helveticus, (Lv) L. vulgaris, (Pw) P. waltl, (Tc) T. cristatus.

Overall, in all species, we see three broad categories of MHC-I genes: (i) putative classical MHC-Ia, expression of which increases in ontogeny, is very high in adults, and they exhibit high polymorphism; (ii) putative nonclassical (MHC-Ib) with low polymorphism, higher relative expression in larval stage and/or tissue-specific expression in adults; and (iii) either nonclassical or nonfunctional, with low expression throughout ontogeny and tissues and, more often than not, low polymorphism. The distribution of these three categories on phylogeny indicates fast turnover and rapid evolution/diversification (Supplemental Fig. S8).
Evolution of the MHC-I gene family in newts
The results presented above indicate a dynamic evolution of MHC-I family in newts, with multiple gene duplications and losses. To obtain a more quantitative picture of these dynamics, we performed gene tree- species tree reconciliation to infer the number of duplications and losses and identify clusters of orthologs. Reconciliation of the MHC-I gene tree with the species tree was performed separately for data sets G and G+T. Each data set produced a single event history—for G, 26 duplications and 10 losses, and for G+T, 91 duplications and 66 losses. All species have experienced lineage-specific duplications and losses, although the exact numerical results in Supplemental Table S10 should be treated with caution because some analyzed sequences may be derived from the same gene. However, we were not able to reliably reconstruct groups of orthologous sequences because, within the same event history, a large number of nodes in the reconciled tree could have been swapped without changing the overall number of events.
Gene conversion may lead to the homogenization of sequences of different members of gene families, resulting in concerted evolution that obliterates the signal of orthology. It is also an important mechanism generating new MHC variants by transferring small sequence segments between alleles and genes. To see whether concerted evolution via interlocus conversion could have contributed to the poor resolution of MHC-I phylogeny and to the tendency for MHC sequences to group by species, we tested for gene conversion, estimated the length of conversion tracts, and checked whether some domains were involved in the process more often than others. For MHC-I-like1 and -like2 which were analyzed together and with all species pooled, the signal of gene conversion was detected only for domain α3, and the tracts were shorter than in MHC-I. In MHC-I, conversion was detected in all species, with a tendency to be more common in α3, which may contribute to the low sequence divergence seen in this domain (Supplemental Fig. S5). We detected both conversion tracts contained within a single domain (i.e., within a single exon), and tracts spanning more than one exon. Because more of the latter were detected when mismatches within tracts were allowed (Supplemental Table S11) and because tracts are expected to be relatively short so that they should not span across long newt introns, we suspect some signal may represent false positives. Nevertheless, the process of gene conversion appears widespread in newt MHC-I and may lead to concerted evolution of MHC genes within species.
Selective pressures across MHC-I phylogeny
We were interested whether MHC-I and two MHC-I-like lineages experienced different selective pressures. In particular, we tested, using RELAX and FUBAR: (i) whether a change in the strength of selection occurred along the long branches separating MHC-I-like1 and MHC-I-like2 from other sequences in MHC-I phylogeny, and (ii) whether the number and location of codons under positive/purifying selection differ among the three categories of genes.
There was evidence of selection relaxation for MHC-I-like1: a highly significant (P-values < 0.001) RELAX test result for the branch leading to MHC-I-like1 cluster (k: 0.356, LRT: 401.7), and for the cluster itself (k: 0.476, LRT: 42.7), as compared to the MHC-I sequences. Accordingly, a direct, codon-based selection test (FUBAR) found only a few codons under positive or negative selection among MHC-I-like1 sequences (three and seven, respectively—posterior probability >0.9). The second divergent clade, MHC-I-like2, was associated with a slight relaxation of selection intensity in the branch leading to the clade (P-value: 0.021, k: 0.857, LRT: 5.3) but not in the cluster itself (P-value: 0.112). A direct selection test found 12 sites under positive and 50 sites under negative selection, which was considerably fewer compared to MHC-I sequences (30 and 122, respectively). Only two of the sites under positive selection overlapped between the two groups (Fig. 5). The selection tests were performed only on the three extracellular domains of MHC-I (α1, α2, α3), because the divergence of the remaining parts (i.e., the connecting, transmembrane, and cytosolic portion of the molecules) rendered the alignment unreliable. Roughly half of the residues in the MHC-I were under negative selection, compared to ∼20% in the case of the MHC-I-like2 sequences. Most of the positively selected sites in MHC-I were located in the α1 and α2 domains (14 and 10, respectively) and largely coincided with or were immediately adjacent to residues forming peptide biding pockets as identified in Xenopus (Fig. 5). MHC-I-like2 sequences had many fewer positions under positive selection, scattered through the extracellular domains (four, six and three in α1, α2, and α3, respectively) (Fig. 5).
The overall concordance between the positively selected sites identified in the sequences forming the major MHC-I clade (Fig. 4A; Supplemental Fig. S1) and the peptide binding pockets identified in Xenopus (Fig. 5) supports the classification of these sequences as the putative MHC-Ia (and perhaps some MHC-Ib). In contrast, markedly relaxed selection in MHC-I-like1 suggests a trajectory leading to pseudogenization, whereas, in the case of MHC-I-like2, change in selective pressure strength and location of codons under positive selection hints to alterations in function, perhaps specialization.
Structural modeling of MHC-I-like sequences
MHC-I-like1 and MHC-I-like2 sequences exhibit characteristic deletions within the α1 and α2 domains (Fig. 5), observed in transcripts and confirmed in the available genomes. MHC-I-like1 molecules lack a signal peptide and have a truncated α1 domain (missing 23 N-terminal amino acids). Lack of a signal peptide likely precludes transport to the endoplasmic reticulum, where MHC-I is folded and loaded with antigens. Thus, this modification alone suggests departure from the classical function. In contrast, MHC-I-like2 molecules possess a short N-terminal sequence resembling a signal peptide, an intact α1 domain, and substantial deletions within the α2 domain. To evaluate the structural impact of these alterations on the peptide-binding domains, we performed in silico modeling using AlphaFold 3. Presumably classical MHC-I proteins, included as references, were modeled with high confidence, exhibiting predicted local distance difference test (pLDDT) scores >90 in the α1–α2 domains. In comparison, the predicted α1–α2 domains of MHC-I-like molecules showed lower confidence, with overall pLDDT scores ranging from >70 to <50 in the most uncertain regions. Most of the MHC-I-like1 models preserved the overall MHC-I fold, including the two antiparallel α-helices, but lacked the first β-strand and loop of the α1 domain (Supplemental Fig. S10A), although one variant showed a closed conformation without open pockets (L. vulgaris conformation 2) (Supplemental Fig. S10A). Structural homology searches identified MHC and MHC-like molecules from other vertebrates as the closest matches, but none of the molecules contained deletions in the discussed region, and thus the role of the deletion remains unknown. We speculate that a molecule with such an extensive deletion in the core would not fold altogether, but this speculation remains to be experimentally tested. In MHC-I-like2 molecules, deletions within the α2 domain affect regions of the peptide-biding groove that, in classical MHC-Ia molecules, anchor the C-termini of bound peptides. Two β-strands forming the groove floor were either missing or significantly shortened, and a substantial portion of the α2 helix (including the conserved R146 and W147 residues) was also absent (Supplemental Fig. S10B). Structural homology searches again primarily returned MHC and MHC-like molecules; these molecules however, lacked the discussed deletion. A notable exception was the viral MHC-I homolog m144 (from murine cytomegalovirus, PDB: 1U58), which shares similar deletions in the α2 domain. Unlike m144, which lacks a well-defined groove, the predicted MHC-I-like2 structures displayed seemingly an open conformation with a spacious, predominantly lipophilic pocket—this may, however, reflect the bias of AlphaFold towards classical MHC structures. Functionally, m144 is a molecular decoy mimicking host MHC-I and expressed by the virus to prolong its survival by preventing natural killer cell activation (Natarajan et al. 2006). Although this does not offer clues to the function of MHC-I-like2 in vertebrate hosts, it does support the plausibility of its folding pattern.
Discussion
The results reported here advance our understanding of MHC evolution in salamanders and have broader relevance beyond this group. We will discuss first the broader implications of our work, that is, the size and structure of the MHC region in giant genomes, coevolution of polymorphic genes within the MHC, and immunity of free-living early developmental stages of aquatic vertebrates. Then, we will present insights more specific to salamanders and conclude with a discussion of the limitations of the study and future prospects.
All studied newt species have a relatively compact core MHC region, which corresponds to the “adaptive MHC” of Veríssimo et al. (2023), although in some species (most notably in T. cristatus), additional MHC-I genes are present at a large distance or on other chromosomes. Nonetheless, the main class I (containing APGs) and class II regions are located next to each other and, unlike in eutherian mammals, are not separated by other genes. This presumably ancestral organization (Ohta and Flajnik 2015; Veríssimo et al. 2023) is thus retained both in salamanders and in frogs (Session et al. 2016; He et al. 2023). Our data do not confirm extensive expansion of the MHC region proposed in salamanders or suggestions that the large-scale MHC architecture in this group may be more similar to that of eutherian mammals (Schloissnig et al. 2021). This earlier work was based on the analysis of synteny and emphasized the genomic location of genes and gene families found in MHC class III or linked to MHC class I in the eutherian mammals. As already demonstrated by Migalska et al. (2025), such an approach may indeed identify large genomic regions, which are, however, devoid of bona fide MHC genes. Thus, describing them as the MHC region is, at best, questionable. The situation is well illustrated by the P. waltl genome, where the genes found in the eutherian MHC region are located in two blocks at the opposite ends of Chr 6, but the MHC genes are found only in a compact region within one of these blocks (Brown et al. 2025; this study). The eutherian gene arrangement is derived, and the position of these genes should not be used to define the MHC region in other taxa. Our results thus demonstrate that the genomic gigantism in distantly related newts is accompanied by a relatively compact MHC region. Whether this holds in other taxa characterized by genomic gigantism is somewhat unclear. In another Urodele species with a huge (∼32 Gb) genome and resolved MHC architecture—the axolotl (Ambystoma mexicanum)—the core MHC spans 13.8 Mb and has analogous organization to that found in newts (Migalska et al. 2025). In lungfish Protopterus annectens, the “adaptive MHC” region encompasses 20 Mb and is also relatively compact, considering the enormous genome size (>40 Gb). However, additional MHC-II genes are located at a considerable distance; therefore, if a broader definition were applied to define it, the MHC region in the lungfish would span over 170 Mb (Veríssimo et al. 2023).
The compact genomic architecture is significant for a major theory positing that the tight linkage between classical MHC-I and APGs facilitates their coevolution, resulting in matching sets of MHC-I and APG alleles grouped in haplotypes (Joly et al. 1998; Kaufman 1999, 2015; Ohta and Flajnik 2015). Close proximity of coevolving polymorphic partners would minimize the formation of unfit MHC-I–APG recombinants. This theory, based on research in chicken and rat, assumes a single dominantly expressed MHC-I gene that coevolves with APGs (Kaufman 1999, 2015). Such a state has been suggested as ancestral for jawed vertebrates, with a “lucky” accident of a genomic rearrangement in mammals that broke the linkage between MHC-I and APGs. This presumably led to the evolution of generalist APGs, which, in turn, allowed duplication of classical MHC-I in mammals and an increased within-individual MHC variation (Kaufman 2018a). Notably, a key exception helped shape this hypothesis: in rats, a translocation of an MHC-Ia gene back into proximity with APGs (and loss of any remaining MHC-Ia genes from their Class I region) appeared to reinstate coevolutionary dynamics (Joly et al. 1998). A major prediction of the coevolution hypothesis—a positive correlation between the MHC-I and APG diversity—was tested in a comparative framework using salamanders (Palomar et al. 2021). Salamanders commonly have more than one highly expressed MHC-I gene (Sammut et al. 1999; Fijarczyk et al. 2018; this study) and highly polymorphic TAP APGs (Fijarczyk et al. 2018; Palomar et al. 2021). Palomar et al. (2021) found a positive correlation between the diversity of MHC-I and TAP genes (but not other APGs), indicating that the coevolution between APGs and more than one classical MHC-I may be possible. The combined evidence from genomics, transcriptomics, and polymorphism presented here sheds further light on this matter. In three out of the four genome assemblies (L. helveticus, L. vulgaris, and P. waltl), the most highly expressed MHC-I gene is indeed the closest to the APG. However, the most highly expressed gene is not always the most polymorphic (although, in all cases, polymorphism is substantial), nor is it necessarily the only highly expressed gene. Lissotriton newts simultaneously have multiple highly polymorphic genes, and two or more genes located in the proximity of APGs are highly expressed—both hallmarks of classical MHC-Ia. Taken together, these results indicate that more than one highly polymorphic and highly expressed MHC-I gene coexists with polymorphic TAP genes (Palomar et al. 2021), speaking against universal coevolutionary constraints forcing a single, dominantly expressed MHC-I gene.
Nonetheless, separation of APGs and MHC-I may indeed both facilitate MHC-I expansion and reduce TAP polymorphism—as exemplified by the exceptional situation found in T. cristatus. There, four highly polymorphic and highly expressed genes are located, not in the core MHC region, where APGs reside, but on another chromosome. In the seven Triturus species studied by Palomar et al. (2021), TAP diversity was among the lowest across salamander species and did not seem to correlate with MHC-I diversity (Supplemental Fig. S11). This is consistent with an emergence of generalist nonpolymorphic TAPs, which supply peptides to the classical proteins encoded by unlinked MHC-I genes. However, to further support this proposition, MHC organization in more Triturus genomes will need to be resolved. Overall, our findings point to a considerable plasticity in the genomic organization of the MHC-I in newts, which may translate into evolutionary flexibility in the interaction between key components of adaptive immunity.
Ontogenetic expression profiles of MHC genes, integrated with tissue-specific expression patterns and polymorphism data, allow (at least approximate) distinction between classical and nonclassical MHC-I genes. We note, however, that the differences in polymorphism among newt MHC-I genes (Fig. 8; Supplemental Table S8) are less clear-cut than in, e.g., humans. For example, the most polymorphic human MHC-Ib gene (HLA-E) has <3% of the alleles found in the least polymorphic MHC-Ia gene (HLA-C), which itself has 77% as many alleles as the most polymorphic gene, HLA-B. All seven newt species examined have some MHC-I genes with high relative expression in larvae but not in adulthood. In the four species with available genomes, these genes also show low polymorphism. Both characteristics support their classification as nonclassical (MHC-Ib) genes, which could play a key role in larval immunity. Such a situation was described in Xenopus frogs, where XNC MHC-Ib restricts a population of unconventional T lymphocytes important in larval immunity (Edholm et al. 2013, 2018). However, unlike Xenopus, putative MHC-Ib genes in newts do not form a divergent gene family but rather emerge repeatedly from classical MHC-I genes. This observation is in line with comparative data indicating that expanded families of putative, divergent MHC-Ib are not widespread among amphibians and that the Xenopus genus might be a notable exception (He et al. 2023). These findings have broad implications for our understanding of developmental immunology in aquatic vertebrates. Such species must fight pathogen assault from early, free-living developmental stages, when constrained by the low number of available lymphocytes and the time needed to develop the diverse adaptive immune repertoire. The use of unconventional lymphocytes that carry T cell receptors of limited diversity could be an efficient solution in the face of such constraints (Edholm et al. 2014b; Robert and Edholm 2014). If such a strategy is indeed widespread, the rapid turnover of MHC-Ib genes, such as we see in newts, would suggest that populations of unconventional, innate-like T cells restricted by MHC-Ib molecules are ancient, and only MHC-Ib undergo rapid turnover, being repeatedly generated from classical MHC-I genes, in response to taxon-specific pathogen pressures. To test this hypothesis, ontogenetic profiling of lymphocyte populations across a wide range of aquatic vertebrates, using single-cell sequencing (Papalexi and Satija 2018; Rubin et al. 2021), will be essential. We also note the evidence for the nonclassical status of genes reported here in newts comes from expression and polymorphism data and does not have a strong, experimental support of Xenopus studies (Goyos et al. 2011; Edholm et al. 2014a). Therefore, complementary, functional studies at broader, phylogenetic scales are needed, although, due to massive technical difficulties, likely are not feasible in the near future.
In addition to the three themes of broad relevance discussed above, our results also shed light on several aspects of the developmental and evolutionary biology of MHC in salamanders. Our data set tracks gene expression through larval ontogeny in newts, providing insights into the development of adaptive immunity in this group of amphibians. Particularly interesting is the contrast between Pleurodeles and the remaining six species, as for P. waltl, substantial MHC expression appears already in larva and increases through metamorphosis. The onset of MHC expression (older larvae, stage 45–47 of Shi and Boucaut 1995) coincides with the moment when thymectomy ceases to abrogate allograft rejection (Fache and Charlemagne 1975), which marks the beginning of adaptive immune capacities. Experimental evidence for other newt species is modest, but some early work suggested the development of adaptive immunity may be generally more gradual in Urodeles, compared to Anurans (for review, see Du Pasquier 1982). Whereas this seems to be the case in P. waltl, poor overall larval MHC expression in other species contradicts this view. Pleurodeles is by far the largest of the taxa studied here, exceeding 75 g body mass, whereas Triturus weighs up to 15 g, Ichthyosaura 7 g, and Lissotriton 5 g. It is therefore possible that species-specific body size (correlating with cell numbers, including lymphoid cells), rather than ontogenetic stage, determine the onset of adaptive immunity. This hypothesis can be tested with more data from salamanders that differ in body size.
Our data reveal several patterns in the evolutionary trajectories of immune genes in newts. The MHC-I gene family is highly dynamic, with abundant lineage-specific duplications, and some degree of concerted evolution through gene conversion resulting in poorly resolved phylogeny and a tendency of genes to cluster by species. Putative MHC-Ib genes originate repeatedly as discussed above, and numerous genes show both low expression and low polymorphism, indicating either nonclassical functions or the loss of functionality. Numerous MHC-I pseudogenes and gene fragments within the MHC region in all species testify to the dynamic birth and death evolution (Nei and Rooney 2005) at the temporal scale comparable to that in mammals or faster (Fortier and Pritchard 2025). The dynamic evolution of MHC-I contrasts with the situation in MHC-II family. Nonclassical DMA and DMB genes exhibit no signs of duplication and low polymorphism within species. There are also generally single classical IIA and IIB genes, with duplication detected only in L. vulgaris. The most remarkable feature of IIA and IIB genes is the presence of deeply divergent allelic lineages in all genera except Pleurodeles, suggesting long-term retention of transspecies polymorphism that may have persisted for >30 million years. In addition, we identified two distinct lineages of MHC-I-like genes, separated by long branches from each other and from other MHC-I sequences. Expression of MHC-I-like genes starts in the early larva and stays similarly low throughout life. MHC-I-like1 is absent in L. helveticus, pseudogenized in T. cristatus, and shows signs of relaxed selection. However, its substantial expression in P. waltl, reaching 30 FPKM, suggests a possible functional role. MHC-like-1 encodes a structurally atypical protein lacking a signal peptide and part of the α1 domain, making it unlikely to function in antigen presentation. In fact, it is unlikely that such protein folds at all in the α1 region. Structural modeling revealed a potentially artifactual cavity in the α1 domain, but we assume it was due to AlphaFold's bias toward canonical MHC-I folds, on which the algorithm is trained. In contrast, MHC-I-like2 genes are present across all species, exhibit modest duplication, rapid evolution, and display a distinct distribution of positively selected sites in the α1 and α2 domains. Existence of a structurally analogous viral protein (m144) supports a functional potential, but the role of MHC-I-like2 is most certainly different from m144 and likely also diverges from the known MHC-I molecules. Definitive conclusions would require experimental structural data and functional assays.
The integration of chromosome-scale genome assemblies and transcriptomic data, particularly PacBio Iso-Seq, has yielded valuable insights into the structure and evolution of the MHC in newts. However, both data types come with limitations. Assembling gigantic newt genomes remains a formidable challenge, making it unlikely that multiple high-quality assemblies will be available soon. Although Iso-Seq data can be obtained from multiple individuals and capture full-length transcripts, it does not provide information about the genomic location of those transcripts—a critical limitation given the complexity and rapid evolution of the newt MHC, including extensive copy number variation. Encouragingly, the relatively moderate length of the MHC region (in relation to the genome size) offers hope that advances in long-fragment DNA capture for long-read sequencing technologies (Iyer et al. 2024) will soon make it feasible to sequence full MHC haplotypes from multiple individuals. Such data are essential for a deeper understanding of the mechanisms driving MHC diversity and for overcoming the inherent limitations of amplicon sequencing routinely used for MHC genotyping. As our results clearly demonstrate, short amplicon data do not allow unambiguous assignment of variants to MHC genes, even when a high-quality reference genome is available. Further data on ontogenetic expression profiles from species that differ in body size are needed to test the hypothesis that body size determines the timing of the onset of adaptive immunity in salamanders. Finally, research on MHC genes should go hand in hand with characterizing the populations of T cells and repertoires of T cell receptors through the ontogeny of diverse salamander species.
Methods
Genome assemblies
We described MHC region(s) in chromosome-scale genome assemblies of P. waltl (obtained from the GenBank database [https://www.ncbi.nlm.nih.gov/genbank/] under accession number GCA_031143425.1) (Brown et al. 2025) and three newt species sequenced and assembled by the Darwin Tree of Life Project: L. helveticus (GCA_964261635.1), L. vulgaris (GCA_964263255.1), and T. cristatus (GCA_964204655.1). The analyses were based on the primary haplotypes under the indicated accession numbers.
RNA-seq
We analyzed transcriptomes of the following newt species: I. alpestris, L. boscai, L. helveticus, L. vulgaris, P. waltl, T. cristatus, and T. marmoratus (Fig. 1A). Separate libraries were prepared for each tissue/body part of each individual (Supplemental Table S4). We did not analyze thymus specifically because, given differences in larval size across species and developmental stages, it was often too small to yield sufficient RNA for library preparation with the chosen methodology. To maintain consistency across samples and taxa, we therefore collected larger tissue fragments. Thymus-derived expression data would be highly informative and represent an important avenue for future studies. Details about samples and laboratory procedures are in Supplemental Tables S2 and S4 and Supplemental Methods.
Full-length transcript sequencing
For P. waltl, we used the available PacBio Iso-Seq data obtained from the spleen RNA of the individual used for genome assembly (Brown et al. 2025). For the remaining six species, Iso-Seq libraries from intestine (a tissue where diverse immune cells are abundant) RNA of a single individual per species were prepared and sequenced on the PacBio Revio platform by Macrogen. Note that, for each of these samples, short-read Illumina data were also generated. Details are in Supplemental Methods.
Identification of MHC transcripts
To identify MHC sequences across newt species, we used axolotl and P. waltl MHC proteins as queries to search Iso-Seq transcriptomes. Candidate transcripts were filtered and clustered to reduce redundancy, and representative sequences were aligned and curated. To capture additional diversity that might be absent from Iso-Seq data, we supplemented these analyses with de novo assemblies of short-read RNA-seq data. Although some of the short-read–derived sequences may be artifactual, they provide complementary evidence for additional MHC diversity. Details are in Supplemental Methods.
The genomic organization of the MHC region
The genomic location of the “core” MHC region and any disparate genome fragments that may contain MHC genes was established by BLASTN searches with cluster representatives obtained as described above. To annotate genes both within these regions and in their vicinity, we adopted two routes; protein-coding genes other than MHC-I and MHC-II were automatically annotated, whereas MHC genes were annotated manually. Details are in Supplemental Methods.
To ease comparative analysis, we classified all the annotated functional genes following categories used by Belov et al. (2006): (i) MHC Class I (putative classical or nonclassical); (ii) MHC Class II (including classical MHC-II, DM, and a non-MHC, yet closely linked with Class II region BRD2 gene); (iii) Antigen Processing Genes (which include PSMB8, PSMB9, TAP1, and TAP2); (iv) Extended class II (from COL11A2 to ZBTB22, if present, including TAPBP); (v) MHC Class III (if present); and (vi) Framework region genes (conserved, non-MHC genes known to be interspersed in the eutherian MHC class I region [Amadou 1999], found in the vicinity of the core MHC region in nonplacental mammals [Belov et al. 2006], and more dispersed in nonmammalian genomes). These categories are well suited to describe noneutherian MHC organization while maintaining clear correspondence to the gold standard of the human MHC nomenclature. For clarity of visualization, we show only synteny-informative Framework genes (i.e., present in more than one species) and omit genes that are not mapping to human MHC.
MHC-I polymorphism
One of the key features distinguishing classical MHC genes from nonclassical is their high polymorphism. To assess polymorphism of the MHC-I genes annotated in the genomes, we used exon 2 amplicon sequencing data from previous studies (Palomar et al. 2021; Gaczorek et al. 2023). Polymorphism was measured as the number of alleles mapped to each MHC-I gene, with corrections applied for alleles mapping to multiple genes, and standardized relative to the most polymorphic gene. Details are in Supplemental Methods. This approach is approximate, as alleles may not always be most similar to the reference sequence of their genes of origin due to the high divergence between alleles and interlocus recombination, the genes of origin may be missing from the reference assembly, and alleles may map equally well to multiple genes. Still, even such an approximate measure should provide meaningful information about polymorphism.
Phylogenies of MHC proteins
Clustering of MHC proteins, following the procedure described above, was repeated to include also the proteins predicted in genomes. When available, the longest proteins predicted in the genome were used as cluster representatives. The sequences formed the G+T dataset which was subject to phylogenetic analyses. Details are in Supplemental Methods.
MHC-I anchor residues
MHC-I anchor residues are defined as amino acids at nine residues important for anchoring the termini of antigenic peptides that are conserved in classical MHC-I of most taxa (Kaufman et al. 1994; Sammut et al. 1999; Almeida et al. 2021). The number of residues with conserved amino acids may be helpful in distinguishing classical and nonclassical MHC-I sequences, and we calculated it for our sequences.
MHC-I and MHC-II expression across tissues and developmental stages
MHC-I and MHC-II expression across tissues and developmental stages was quantified by mapping RNA-seq data to genome assemblies (where available) and to representative full-length MHC coding sequences (details are in Supplemental Methods). Expression was calculated as fragments per kilobase of transcript length per million reads mapped (FPKM), and relative expression was defined as the fraction of total MHC-I expression attributable to a given sequence in each library. For cases requiring a single expression value (e.g., for visualization), we used the maximum relative expression across all RNA-seq libraries from postmetamorphic individuals of a species.
Gene tree–species tree reconciliation
Gene tree-species tree reconciliation was performed in Notung (Durand et al. 2006) to infer the number of duplications and gene losses and identify clusters of orthologs in MHC-I. Details are in Supplemental Methods.
Gene conversion
Gene conversion was tested using coding DNA sequences of the G+T data set after removing sequences assembled by Trinity (Grabherr et al. 2011) from short RNA-seq reads, as some of them could have been computationally generated chimeras that may produce a false recombination signal. Details are in Supplemental Methods.
Selective pressures acting on MHC-I and MHC-I-like genes
Selective pressures acting on MHC-I and MHC-I-like genes were investigated using HyPhy (Kosakovsky Pond et al. 2020), focusing on whether selection pressures differed along major phylogenetic branches and between clades, and on the identification of codons under positive or purifying selection. Details are in Supplemental Methods.
Structural modeling
Structural modeling of MHC-I-like molecules and representative classical MHC-I molecules from each species, together with β2m, was performed using AlphaFold3 (Abramson et al. 2024). The resulting models were inspected for structural features and compared to known proteins to assess homology. Details are in Supplemental Methods.
Data access
Raw sequence data generated in this study have been submitted to the European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena/browser/home) under accession number PRJEB90989. Sequence alignments and de novo annotation of MHC genes are available as Supplemental Data.
Competing interest statement
The authors declare no competing interests.
Acknowledgments
This study was supported by the Polish National Science Centre OPUS Grant No. UMO-2021/43/B/NZ8/00979 to W.B. The authors thank Feliks Ohlaszeny for his help in laboratory analyses.
Author contributions: W.B. designed and supervised the study; K.D., M.Mi., G.P., and W.B. collected the samples; M.H.Y. provided Pleurodeles adults; K.D. and M.Ma. collected the data; W.B., M.Mi., K.D., and G.D. performed data analyses; and W.B. and M.Mi. drafted and revised the manuscript. All authors reviewed and contributed to the writing of the final manuscript.
Footnotes
[1] Supplementary material [Supplemental material is available for this article.]
[2] Article published online before print. Article, supplemental material, and publication date are at https://www.genome.org/cgi/doi/10.1101/gr.281127.125.
References
- ↵Abramson J, Adler J, Dunger J, Evans R, Green T, Pritzel A, Ronneberger O, Willmore L, Ballard AJ, Bambrick J, 2024. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630: 493–500. 10.1038/s41586-024-07487-w
- ↵Adams EJ, Luoma AM. 2013. The adaptable major histocompatibility complex (MHC) fold: structure and function of nonclassical and MHC class I–like molecules. Annu Rev Immun 31: 529–561. 10.1146/annurev-immunol-032712-095912
- ↵Almeida T, Ohta Y, Gaigher A, Muñoz-Mérida A, Neves F, Castro LFC, Machado AM, Esteves PJ, Veríssimo A, Flajnik MF. 2021. A highly complex, MHC-linked, 350 million-year-old shark nonclassical class I lineage. J Immun 207: 824–836. 10.4049/jimmunol.2000851
- ↵Amadou C. 1999. Evolution of the Mhc class I region: the framework hypothesis. Immunogenetics 49: 362–367. 10.1007/s002510050507
- ↵Belov K, Deakin JE, Papenfuss AT, Baker ML, Melman SD, Siddle HV, Gouin N, Goode DL, Sargeant TJ, Robinson MD, 2006. Reconstructing an ancestral mammalian immune supercomplex from a marsupial major histocompatibility complex. PLoS Biol 4: e46. 10.1371/journal.pbio.0040046
- ↵Bentley BP, Carrasco-Valenzuela T, Ramos EKS, Pawar H, Souza Arantes L, Alexander A, Banerjee SM, Masterson P, Kuhlwilm M, Pippel M, 2023. Divergent sensory and immune gene evolution in sea turtles with contrasting demographic and life histories. Proc Natl Acad Sci 120: e2201076120. 10.1073/pnas.2201076120
- ↵Brown T, Mishra K, Elewa A, Iarovenko S, Subramanian E, Araus AJ, Petzold A, Fromm B, Friedländer MR, Rikk L, 2025. Chromosome-scale genome assembly reveals how repeat elements shape non-coding RNA landscapes active during newt limb regeneration. Cell Genom 5: 100761. 10.1016/j.xgen.2025.100761
- ↵Card DC, Van Camp AG, Santonastaso T, Jensen-Seaman MI, Anthony NM, Edwards SV. 2022. Structure and evolution of the squamate major histocompatibility complex as revealed by two Anolis lizard genomes. Front Genet 13: 979746. 10.3389/fgene.2022.979746
- ↵Cohen N. 1980. Salamanders and the evolution of the major histocompatibility complex. In Contemporary topics in immunobiology (ed. Marchalonis JJ, Cohen N), pp. 109–139. Springer US, Boston.
- ↵Du Pasquier L. 1982. Ontogeny of immunological functions in amphibians. In Phylogeny and ontogeny (ed. Cohen N, Sigel MM), pp. 633–657. Springer US, Boston.
- ↵Durand D, Halldórsson BV, Vernot B. 2006. A hybrid micro–macroevolutionary approach to gene tree reconstruction. J Comput Biol 13: 320–335. 10.1089/cmb.2006.13.320
- ↵Edholm E-S, Saez L-MA, Gill AL, Gill SR, Grayfer L, Haynes N, Myers JR, Robert J. 2013. Nonclassical MHC class I-dependent invariant T cells are evolutionarily conserved and prominent from early development in amphibians. Proc Natl Acad Sci 110: 14342–14347. 10.1073/pnas.1309840110
- ↵Edholm E-S, Goyos A, Taran J, Andino FDJ, Ohta Y, Robert J. 2014a. Unusual evolutionary conservation and further species-specific adaptations of a large family of nonclassical MHC class Ib genes across different degrees of genome ploidy in the amphibian subfamily Xenopodinae. Immunogenetics 66: 411–426. 10.1007/s00251-014-0774-5
- ↵Edholm E-S, Grayfer L, Robert J. 2014b. Evolution of nonclassical MHC-dependent invariant T cells. Cell Mol Life Sci 71: 4763–4780. 10.1007/s00018-014-1701-5
- ↵Edholm E-S, Banach M, Rhoo KH, Pavelka MS, Robert J. 2018. Distinct MHC class I-like interacting invariant T cell lineage at the forefront of mycobacterial immunity uncovered in Xenopus. Proc Natl Acad Sci 115: E4023–E4031. 10.1073/pnas.1722129115
- ↵Fache B, Charlemagne J. 1975. Influence on allograft rejection of thymectomy at different stages of larval development in urodele amphibian Pleurodeles waltlii Michah. (Salamandridae). Eur J Immunol 5: 155–157. 10.1002/eji.1830050215
- ↵Fijarczyk A, Dudek K, Niedzicka M, Babik W. 2018. Balancing selection and introgression of newt immune-response genes. Proc Roy Soc B 285: 20180819. 10.1098/rspb.2018.0819
- ↵Flajnik MF, Kasahara M. 2010. Origin and evolution of the adaptive immune system: genetic events and selective pressures. Nat Rev Genet 11: 47–59. 10.1038/nrg2703
- ↵Fortier AL, Pritchard JK. 2025. The primate Major Histocompatibility Complex as a case study of gene family evolution. eLife 14: RP103545. 10.7554/eLife.103545
- ↵France J, Babik W, Cvijanović M, Dudek K, Ivanović A, Vučić T, Wielstra B. 2025. Identification of Y-chromosome turnover in newts fails to support a sex chromosome origin for the Triturus balanced lethal system. Genome Biol Evol 17: evaf155. 10.1101/2024.11.04.621952
- ↵Gaczorek TS, Marszałek M, Dudek K, Arntzen JW, Wielstra B, Babik W. 2023. Interspecific introgression of MHC genes in Triturus newts: evidence from multiple contact zones. Mol Ecol 32: 867–880. 10.1111/mec.16804
- ↵Goyos A, Sowa J, Ohta Y, Robert J. 2011. Remarkable conservation of distinct nonclassical MHC class I lineages in divergent amphibian species. J Immunol 186: 372–381. 10.4049/jimmunol.1001467
- ↵Grabherr MG, Haas BJ, Yassour M, Levin JZ, Thompson DA, Amit I, Adiconis X, Fan L, Raychowdhury R, Zeng Q, 2011. Full-length transcriptome assembly from RNA-seq data without a reference genome. Nat Biotech 29: 644–652. 10.1038/nbt.1883
- ↵Grimholt U, Tsukamoto K, Azuma T, Leong J, Koop BF, Dijkstra JM. 2015. A comprehensive analysis of teleost MHC class I sequences. BMC Evol Biol 15: 32. 10.1186/s12862-015-0309-1
- ↵He K, Zhu Y, Yang S-C, Ye Q, Fang S-G, Wan Q-H. 2022. Major histocompatibility complex genomic investigation of endangered Chinese alligator provides insights into the evolution of tetrapod major histocompatibility complex and survival of critically bottlenecked species. Front Ecol Evol 10: 1078058. 10.3389/fevo.2022.1078058
- ↵He K, Babik W, Majda M, Minias P. 2023. MHC architecture in amphibians—ancestral reconstruction, gene rearrangements, and duplication patterns. Genome Biol Evol 15: evad079. 10.1093/gbe/evad079
- ↵Horton R, Wilming L, Rand V, Lovering RC, Bruford EA, Khodiyar VK, Lush MJ, Povey S, Talbot CC, Wright MW, 2004. Gene map of the extended human MHC. Nat Rev Genet 5: 889–899. 10.1038/nrg1489
- ↵Hughes AL, Nei M. 1988. Pattern of nucleotide substitution at major histocompatibility complex class I loci reveals overdominant selection. Nature 335: 167–170. 10.1038/335167a0
- ↵Hughes AL, Nei M. 1992. Maintenance of MHC polymorphism. Nature 355: 402–403. 10.1038/355402b0
- ↵Iyer SV, Goodwin S, McCombie WR. 2024. Leveraging the power of long reads for targeted sequencing. Genome Res 34: 1701–1718. 10.1101/gr.279168.124
- ↵Joly E, Le Rolle AF, Gonzélez AL, Mehling B, Stevens J, Coadwell WJ, Hünig T, Howard JC, Butcher GW. 1998. Co-evolution of rat TAP transporters and MHC class I RT1-A molecules. Curr Biol 8: 169–180. 10.1016/S0960-9822(98)70065-X
- ↵Kaufman J. 1999. Co-evolving genes in MHC haplotypes: the “rule” for nonmammalian vertebrates? Immunogenetics 50: 228–236. 10.1007/s002510050597
- ↵Kaufman J. 2015. Co-evolution with chicken class I genes. Immun Rev 267: 56–71. 10.1111/imr.12321
- ↵Kaufman J. 2018a. Generalists and specialists: a new view of how MHC class I molecules fight infectious pathogens. Trends Immun 39: 367–379. 10.1016/j.it.2018.01.001
- ↵Kaufman J. 2018b. Unfinished business: evolution of the MHC and the adaptive immune system of jawed vertebrates. Annu Rev Immun 36: 383–409. 10.1146/annurev-immunol-051116-052450
- ↵Kaufman J, Volk H. 1994. The salamander immune system: right and wrong. Axolotl Newsletter 23: 7–23. https://ambystoma.uky.edu/genetic-stock-center/newsletters/Older_archive/Issues-13-23/archive/issue23/07-23kaufman%20volk.pdf.
- ↵Kaufman J, Salomonsen J, Flajnik M. 1994. Evolutionary conservation of MHC class I and class II molecules—different yet the same. Sem Immun 6: 411–424. 10.1006/smim.1994.1050
- ↵Kaufman J, Völk H, Wallny HJ. 1995. A “minimal essential Mhc” and an “unrecognized Mhc”: two extremes in selection for polymorphism. Immun Rev 143: 63–88. 10.1111/j.1600-065X.1995.tb00670.x
- ↵Kaufman J, Milne S, Göbel TW, Walker BA, Jacob JP, Auffray C, Zoorob R, Beck S. 1999. The chicken B locus is a minimal essential major histocompatibility complex. Nature 401: 923–925. 10.1038/44856
- ↵Klein J. 1986. The natural history of the major histocompatibility complex. Wiley & Sons, New York.
- ↵Kosakovsky Pond SL, Poon AF, Velazquez R, Weaver S, Hepler NL, Murrell B, Shank SD, Magalis BR, Bouvier D, Nekrutenko A, 2020. Hyphy 2.5—a customizable platform for evolutionary hypothesis testing using phylogenies. Mol Biol Evol 37: 295–299. 10.1093/molbev/msz197
- ↵Martin R, Kaufman J. 2022. Some thoughts about what non-mammalian jawed vertebrates are telling us about antigen processing and peptide loading of MHC molecules. Curr Opin Immunol 77: 102218. 10.1016/j.coi.2022.102218
- ↵Mayassi T, Barreiro LB, Rossjohn J, Jabri B. 2021. A multilayered immune system through the lens of unconventional T cells. Nature 595: 501–510. 10.1038/s41586-021-03578-0
- ↵Migalska M, Sebastian A, Radwan J. 2019. Major histocompatibility complex class I diversity limits the repertoire of T cell receptors. Proc Natl Acad Sci 116: 5021–5026. 10.1073/pnas.1807864116
- ↵Migalska M, Fluder K, Dudek K, Babik W. 2025. Compact genomic architecture of the axolotl MHC region: setting the record straight. Immunogenetics 77: 33. 10.1007/s00251-025-01391-x
- ↵Minias P, Pikus E, Whittingham LA, Dunn PO. 2019. Evolution of copy number at the MHC varies across the avian tree of life. Genome Biol Evol 11: 17–28. 10.1093/gbe/evy253
- ↵Minias P, Palomar G, Dudek K, Babik W. 2022. Salamanders reveal novel trajectories of amphibian MHC evolution. Evolution (N Y) 76: 2436–2449. 10.1111/evo.14601
- ↵Murphy K, Weaver C. 2016. Janeway's immunobiology. Garland Science, New York.
- ↵Natarajan K, Hicks A, Mans J, Robinson H, Guan R, Mariuzza RA, Margulies DH. 2006. Crystal structure of the murine cytomegalovirus MHC-I homolog m144. J Mol Biol 358: 157–171. 10.1016/j.jmb.2006.01.068
- ↵Nei M, Rooney AP. 2005. Concerted and birth-and-death evolution of multigene families. Annu Rev Genet 39: 121–152. 10.1146/annurev.genet.39.073003.112240
- ↵Nowoshilow S, Schloissnig S, Fei J-F, Dahl A, Pang AW, Pippel M, Winkler S, Hastie AR, Young G, Roscito JG, 2018. The axolotl genome and the evolution of key tissue formation regulators. Nature 554: 50–55. 10.1038/nature25458
- ↵Ohta Y, Flajnik MF. 2015. Coevolution of MHC genes (LMP/TAP/class Ia, NKT-class Ib, NKp30-B7H6): lessons from cold-blooded vertebrates. Immun Rev 267: 6–15. 10.1111/imr.12324
- ↵Ohta Y, Goetz W, Hossain MZ, Nonaka M, Flajnik MF. 2006. Ancestral organization of the MHC revealed in the amphibian Xenopus. J Immun 176: 3674–3685. 10.4049/jimmunol.176.6.3674
- ↵Ohta Y, Kasahara M, O'Connor TD, Flajnik MF. 2019. Inferring the “primordial immune complex”: origins of MHC class I and antigen receptors revealed by comparative genomics. J Immun 203: 1882–1896. 10.4049/jimmunol.1900597
- ↵Okamura K, Dijkstra JM, Tsukamoto K, Grimholt U, Wiegertjes GF, Kondow A, Yamaguchi H, Hashimoto K. 2021. Discovery of an ancient MHC category with both class I and class II features. Proc Natl Acad Sci 118: e2108104118. 10.1073/pnas.2108104118
- ↵Palomar G, Dudek K, Migalska M, Arntzen JW, Ficetola GF, Jelić D, Jockusch E, Martínez-Solano I, Matsunami M, Shaffer HB, 2021. Coevolution between MHC class I and antigen-processing genes in salamanders. Mol Biol Evol 38: 5092–5106. 10.1093/molbev/msab237
- ↵Papalexi E, Satija R. 2018. Single-cell RNA sequencing to explore immune cell heterogeneity. Nat Rev Immun 18: 35–45. 10.1038/nri.2017.76
- ↵Pasquier LD, Schwager J, Flajnik MF. 1989. The immune system of Xenopus. Annu Rev Immunol 7: 251–275. 10.1146/annurev.iy.07.040189.001343
- ↵Pishesha N, Harmand TJ, Ploegh HL. 2022. A guide to antigen processing and presentation. Nat Rev Immun 22: 751–764. 10.1038/s41577-022-00707-2
- ↵Radwan J, Babik W, Kaufman J, Lenz TL, Winternitz J. 2020. Advances in the evolutionary understanding of MHC polymorphism. Trends Genet 36: 298–311. 10.1016/j.tig.2020.01.008
- ↵Robert J, Edholm E-S. 2014. A prominent role for invariant T cells in the amphibian Xenopus laevis tadpoles. Immunogenetics 66: 513–523. 10.1007/s00251-014-0781-6
- ↵Rubin SA, Baron CS, Corbin AF, Yang S, Zon LI. 2021. Single-cell transcriptional profiling of zebrafish hematopoiesis offers insight into early lymphocyte development and reveals novel immune cell populations. Blood 138: 4294. 10.1182/blood-2021-146374
- ↵Sammut B, Du Pasquier L, Ducoroy P, Laurens V, Marcuz A, Tournefier A. 1999. Axolotl MHC architecture and polymorphism. Eur J Immunol 29: 2897–2907. 10.1002/(SICI)1521-4141(199909)29:09<2897::AID-IMMU2897>3.0.CO;2-2
- ↵Schloissnig S, Kawaguchi A, Nowoshilow S, Falcon F, Otsuki L, Tardivo P, Timoshevskaya N, Keinath MC, Smith JJ, Voss SR, 2021. The giant axolotl genome uncovers the evolution, scaling, and transcriptional control of complex gene loci. Proc Natl Acad Sci 118: e2017176118. 10.1073/pnas.2017176118
- ↵Session AM, Uno Y, Kwon T, Chapman JA, Toyoda A, Takahashi S, Fukui A, Hikosaka A, Suzuki A, Kondo M, 2016. Genome evolution in the allotetraploid frog Xenopus laevis. Nature 538: 336–343. 10.1038/nature19840
- ↵Shi DL, Boucaut JC. 1995. The chronological development of the urodele amphibian Pleurodeles waltl (Michah). Int J Dev Biol 39: 427–441.
- ↵Shiina T, Hosomichi K, Inoko H, Kulski JK. 2009. The HLA genomic loci map: expression, interaction, diversity and disease. J Human Genet 54: 15–39. 10.1038/jhg.2008.5
- ↵Shiina T, Blancher A, Inoko H, Kulski JK. 2017. Comparative genomics of the human, macaque and mouse major histocompatibility complex. Immunology 150: 127–138. 10.1111/imm.12624
- ↵Smith JJ, Timoshevskaya N, Timoshevskiy VA, Keinath MC, Hardy D, Voss SR. 2019. A chromosome-scale assembly of the axolotl genome. Genome Res 29: 317–324. 10.1101/gr.241901.118
- ↵Stewart AA, Wiens JJ. 2025. A time-calibrated salamander phylogeny including 765 species and 503 genes. Mol Phylogenet Evol 204: 108272. 10.1016/j.ympev.2024.108272
- ↵Veríssimo A, Castro LFC, Muñoz-Mérida A, Almeida T, Gaigher A, Neves F, Flajnik MF, Ohta Y. 2023. An ancestral major histocompatibility complex organization in cartilaginous fish: reconstructing MHC origin and evolution. Mol Biol Evol 40: msad262. 10.1093/molbev/msad262
- ↵Zhou Y, Shearwin-Whyatt L, Li J, Song Z, Hayakawa T, Stevens D, Fenelon JC, Peel E, Cheng Y, Pajpach F, 2021. Platypus and echidna genomes reveal mammalian biology and evolution. Nature 592: 756–762. 10.1038/s41586-020-03039-0