Preprint
Review

This version is not peer-reviewed.

Spliceosomal Ribozyme Assembly

Submitted:

16 September 2026

Posted:

21 September 2026

You are already at the latest version

Abstract
The spliceosome became very complex in eukaryogenesis. The Group IIA intron progenitor was associated with a single protein, homologous to spliceosomal Prp8, but the LECA spliceosome included ~140 proteins. The acquisition of proteins was a neutral process, providing a pool of factors for the development of a coordinated assembly process, where proteins act as scaffold and chaperones, supporting RNA moieties. The RNA component of spliceosomal complexes is tiny. Structural studies offer us snapshots of protein re-arrangements remodelling the RNA. The spliceosome is commonly described as a ‘protein directed ribozyme’. What does this mean? Just how much control ribozymes can delegate to proteins? Spliceosomal ribozymes never lost their primary function of guiding catalysis by RNA base-pairing. To help with alternative splice site choices and to enforce precision, the spliceosome recruited another two small RNAs, U1 and U4, and still employs base-pairing. We discuss RNA structures central in spliceosomal and Group IIA intron ribozymes. Spliceosomal introns preserve protosplice site repeats CAG|GU at 5’ss and 3’ss that dictate a strict order of ribozyme folding. The demarcation of the 5’ss must involve the 3’ss in the downstream repeat. The distinct 5’-3’ss pair of spliceosomal introns serves to reconstruct the correct splice junction between the two repeats, preventing the intron ends from binding U5 snRNA Loop1. Although modern Group IIA introns never splice within repeats, structural and biochemical studies confirm 3’ss involvement at pre-catalytic stage. The exact configuration of the 5’-3’ss pair is different in Group IIA introns, but the parallel strands orientation is conserved. Our updated U5 model, that includes the 5’-3’ss pair, shows that the pre-mRNA strand flipping to achieve the local parallel orientation occurs after the short 3’exon duplex. This asymmetric 5’ and 3’exon binding with the recognition loop is shared with Group IIA introns. The exons are aligned for ligation on U5 Loop1, guided by Watson-Crick pairs as we have previously concluded based on positional dependencies at human splice sites. Our U5 model shows that exon duplexes of stacked pairs are demarcated from the 5’-3’ss pair by a gap, which is how the ribozyme structurally defines the cleavage sites.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Part 1 of this series (Artemyeva-Isman, 2026a) examined the structure and evolution of human intron ribozymes. Fragmentation of eukaryotic intron progenitor into universal trans-acting snRNAs and erosion of cis ribozyme motifs were largely neutral processes, that initially contributed to the burden of early eukaryotic innovations. However, ribozyme assembly in trans with ‘weak’ splice sites was the precondition for alternative splicing leading to radiation of eukaryotic functions that enabled the transition to multicellularity. Alternative splicing flexibility must be combined with precision of mRNA production for a cell to survive. How are flexibility and precision balanced by spliceosomal ribozymes? The spliceosome became very complex in eukaryogenesis: intron progenitor was associated with a single protein, the precursor of spliceosomal largest protein Prp8, but the fully fledged LECA spliceosome included some 140 proteins. Again, acquisition of spliceosomal complexity was a neutral process and initially reduced splicing efficiency, but it provided a pool of factors for the development of a highly coordinated assembly process, where proteins act as scaffold and chaperones, supporting RNA moieties and forming distinct RNP complexes. The RNA component at the core of spliceosomal complexes is relatively small. Structural studies offer us snapshots of protein re-arrangements in the spliceosome that remodel the RNA. The spliceosome is described as a ‘protein directed ribozyme’. What does this mean? Just how much control ribozyme can delegate to proteins? Spliceosomal ribozymes never lost their primary function of guiding catalysis by RNA base-pairing. To help with alternative splice site choices and to enforce precision spliceosome acquired two other small RNAs and still employs base-pairing. Splice sites are integral parts of the ribozyme: their variability, which is necessary for alternative choices, is controlled by the need to preserve ribozyme structure. This is spectacularly demonstrated by the preservation of protosplice site repeats CAG|GU that persist at both 5’ and 3’ splice sites of human introns. Correct splicing within the repeats imposes a strict order on ribozyme folding: demarcation of the 5’ss requires the involvement of the 3’ss in the downstream repeat. We will go through spliceosomal ribozyme assembly step-by-step to examine how flexibility and precision are maintained and how ribozyme folding assisted by proteins directs splicing catalysis. Focusing on the role of the conserved pairing between the guanines of the 5’ and 3’ss, I offer a predicted structure of the pre-catalytic stage U5 snRNA binding the exons.
Figures are numbered throughout this series of linked papers, Part 1 (Artemyeva-Isman, 2026a): Figures 1–4; Part 2 (this paper): Figures 5–7.

2. Early Spliceosome Complexes: Splice Site Choice

2.1. 5’ss Binding by U1 snRNA

The free 5’ end of U1 snRNA recognises the exon/intron boundary and can base-pair with the 3 last positions of the exon and the 8 first positions of the intron (Figure 5,Figure 6A). On average human 5’ splice sites can form only 6-7 Watson-Crick pairs with U1 (Carmel et al., 2004), which is ideal for splicing kinetics, as increasing complementarity from 7 to 11 pairs slows down U1 disassociation and the transition to the pre-catalytic spliceosome, complex B (Lund and Kjems, 2002). U1 snRNA is abundantly expressed and binds multiple alternative and ‘cryptic’ sites of which only one is selected for splicing depending on U1 base-pairing and relative position from U2 snRNP that binds the 3’ss (Hodson et al., 2012; see below 2.4 5’ and 3’ss coupling). Pre-selection of 5’ splice sites and establishing contacts with U2 snRNP are not the only functions of U1 snRNP: it regulates transcription by defining direction at otherwise bi-directional promoters (Almada et al., 2013), prevents premature termination at cryptic intronic polyadenylation signals PAS (Chiu et al., 2018), and increases the elongation rate for synthesis of long transcripts (Mimoso and Adelman, 2023). Notably, mutations in the free 5’end of U1 snRNA 3A C and 7A G are frequently linked to cancers (Bousquets-Muñoz et al., 2022).
Despite its role in searching for possible sites, it was apparent very early on, that U1 base-pairing with pre-mRNA is not responsible for 5’ss definition. Cohen et al., 1994 discovered that U1 can work ‘from a distance’: U1 engineered to bind in the vicinity of its usual binding site at the exon-intron boundary can be as effective for correcting 5’ss mutations, as U1 made complementary to the mutated 5’ss itself (the latter established by Zhuang and Weiner, 1986). Hwang and Cohen, 1996 proceeded to experiment with U6 base-pairing to 5’ss and were the first to suggest that the 5’ cleavage site is defined relative to U6 binding, while U1 merely serves to bring U6 into the vicinity of the 5’ss. In practice, after Fernandez Alanis et al., 2012 U1 variants designed to bind downstream from the mutant 5’ss are widely used to boost splicing in human cells and mice (Donadon et al., 2019; Romano et al., 2022; Jüschke et al., 2021, see Part 3, Artemyeva-Isman, 2026c). It is an accepted view that 5’ss definition is left until the pre-catalytic complex B, when U6 and U5 snRNAs replace U1 (Figure 5,6A).
Recalling the protosplice site repeat CAG|GU at the 3’ss, we come to an unexpected conclusion: U1 can bind at least some 3’ss depending on the downstream sequence of the 3’exon. U1 engineered to bind 20-25nt across 3’ss efficiently suppresses exon inclusion (Hatch et al., 2022; Covello et al., 2022; see Part 3). It remains to be seen if this is true for exons starting with a sequence like the start of the intron |GUAAGUAU and complementary to the wild type U1 5’end.
U1 is entirely dispensable for splicing catalysis and evolved to support the multiple choice of alternative 5’ss (and possibly, 3’ss as we conclude above). Indeed, monocellular eukaryotes have little need for alternative splicing and U1. Famously, monocellular red algae Cyanidioschyzon merolae lost U1 and all its protein co-factor genes altogether (Stark et al., 2015). Splicing catalysis does not require U1 base-pairing with 5’ss, as it proceeds despite the deletion of U1 5’end in human nuclear extracts (Lund and Kjems, 2002). Furthermore, U1 depletion in HeLa cells can be compensated by SR proteins and splicing can be boosted by increasing U6 snRNA complementarity with the start of the intron (Crispino and Sharp, 1995). It appears that U1-independent splicing has a distinct role in human cells, as at least a fraction of introns relies on this mechanism for alternative splicing (Fukumura and Inoue, 2009).
Figure 5. Spliceosomal ribozyme assembly. Each snRNA arrives in complex with protein co-factors – coloured bobbles represent snRNPs. Early complexes: Initially (complex E) several 5’ss are selected by U1 snRNP and possible 3’ss-PPT-BP combinations are identified by U2AF35/65 and SF1 proteins. U2AF65 curves the PPT with a sharp kink, juxtaposing BP and the 3’ss. The first structure of the ribozyme to be formed is the BP Helix: SF3A1 displaces SF1 and recruits U2 snRNP, U2 snRNA searches for homology by branchpoint-interacting stem-loop (BSL) which establishes a ‘toehold’ immediately upstream of the BP A and the BP Helix is formed in place of the BSL stem-loop by RNA strand invasion mechanism (Figure 6C). Based on reviewed structural, biochemical, mutation and phylogenetic evidence (Sickmier et al., 2006; Biancon et al., 2022; Yoshida et al., 2020; Kent et al., 2003; Irimia and Roy, 2008; Corrionero et al., 2011) we suggest, that U2 snRNA bridges across the PPT and fixes BP Helix by the conserved G=C pair with intron position -3. SF3B1 of the large U2-associated SF3A/B protein complex supports U2 snRNA base-pairing with variant BP sequences. Transition to complex A requires mutual stabilisation of U1 and U2 snRNPs: 5’ss and 3’ss are matched which is mediated by U2-associated SF3A1 binding to U1 stem-loop SL4, and UAP56 or SRSF1 that bridge U1 SL3 and U2AF65 or 35 respectively. At this stage cryptic sites are rejected, and the choice of alternative sites can be influenced by RBPs. By contrast, minor introns feature well-conserved splice sites, that do not admit for alternatives and in the minor spliceosome assembly U11 and U12 come together as di-snRNP, skipping the initial pre-selection and ensuring stable binding to both 5’ss and 3’ss at once. Precatalytic spliceosomes: U5•U6/U4 tri-snRNP joins the Pre-B complex and U2 5’ end pairs with U6 forming U6/U2 Helix II (Figure 6B). U6 ACAGAGA sequence is looped out between U6/U4 stem III and U4 quasi-pseudoknot supported by RBM42 bound to U4 opposite the loop (Figure 6A). Transition to B complex involves U1 snRNP disassociation from the 5’ss. Prp28 does not unwind the RNA duplex but instead destabilises it by interfering with the supporting protein U1C (Chen et al., 2001). U6 ACAGAGA loop binds the start of the intron, which leads to the disruption of the quasi-pseudoknot, removal of RBM42 and unwinding of U4/U6 stem III. Brr2 helicase loads to its binding site on U4 in place of stem III. All structural studies agree that at this stage 5’exon binds U5 Loop 1. Here we point out protosplice site repeats CAG|GU, conserved at both 5’ss and 3’ss. A robust mechanism is required to choose the correct junction of exons within repeats. The intron ends G+1••G-1 pair provides an essential link that allows the ribozyme to reposition intermediates between the two steps of splicing and to function in this way, it needs to be formed at pre-catalytic stage. Here we argue that the early involvement of the 5’ss-3’ss pair also provides a solution to the problem of operating with repeats. Splicesome structures do not include 3’ss at the pre-catalytic stage, here we follow the pre-catalytic structure of the Ll.LtrB Group IIA intron, that shows both 5’ss and 3’ss in proximity and the 3’exon aligned with the recognition loop (Liu et al., 2020). We showed previously (Artemyeva-Isman and Porter, 2021) that human U5 and U6 mutually compensate for variations in their binding sites: extensive U5 base-pairing with the 5’exon can support U6 binding to divergent introns and vice versa. This mechanism was cited and confirmed by later studies of different eukaryotes (Ishigami et al., 2021; Parker et al., 2022). We suggested the same mechanism for U5 and U2 at the 3’ss. The importance of specific U5 base-pairing with the 3’exon explained the effect of human mutations on splicing (Artemyeva-Isman and Porter, 2021) and transpired in later phylogenetic and mutation studies (Olthof et al., 2024; Negi et al., 2025; Nava et al., 2025). It is clear, that not only the U6 pairing with the intron, but overall stability of U6, U5 and U2 base-pairing with the pre-mRNA is the fidelity check at the pre-catalytic stage. Transition to the catalytic spliceosome starts with Brr2 helicase translocating on U4 in the 3’ to 5’ direction unwinding U4/U6 stems I and II. The catalytic triple helix of U2/U6 and U6 ISL are formed to coordinate the two Mg2+ ions. NTC/NTR (Nineteen complex, Prp19-associated proteins) join the complex to support the catalytic site conformation (Bact complex) and finally SF3A/B complex disassociates: SF3B1 quits the BP Helix. The bulged adenosine is clumped from Watson-Crick and Hoogsteen edges by the conserved U in position BP-2 and the invariant 3’ss A-2, respectively (Figure 2A). In this configuration the sugar edge of the BP A is facing outwards and ready for 2’O to attack 5’ss (complex B*, here designated C Step 1). Branching triggers a rotation of the BP Helix with the lariat intermediate, the 5’ss-3’ss pair serves to tow the 3’ss and the 3’ exon into the catalytic core. This conformational change is eased by the removal of the 1st step splicing factors. At the 2nd step 5’exon 3’OH attacks the 3’ss, the exons ligate and the intron lariat is freed, C Step 2 complex has the addition of the 2nd lot of splicing factors for conformational support. Next complex P (post-catalytic) still has the exons paired with U5 Loop1, but Prp22 facilitates mRNA release and the complex transitions to intron-lariat spliceosome ILS, that still has the catalytic site, U2 and U6 paired with the intron lariat and U5 snRNA with Loop1 unpaired and in heptaloop closed conformation. ILS is homologous to mobile Group IIA intron particle which can invade new loci by reverse splicing. SF1- Splicing Factor 1; U2AF- U2 Auxiliary Factor 65/35 heterodimer; Prp proteins – Pre-mRNA Processing; Brr2helicase (Bad Response to Refrigeration in yeast mutants) is U5-associated, binding site is indicated on U4 snRNA; 1st Step splicing factors: Cwc25, Yju2, Isy1; 2nd Step factors: FAM32A, Cactin, NKAP and Slu7 (Wilkinson et al., 2017).
Figure 5. Spliceosomal ribozyme assembly. Each snRNA arrives in complex with protein co-factors – coloured bobbles represent snRNPs. Early complexes: Initially (complex E) several 5’ss are selected by U1 snRNP and possible 3’ss-PPT-BP combinations are identified by U2AF35/65 and SF1 proteins. U2AF65 curves the PPT with a sharp kink, juxtaposing BP and the 3’ss. The first structure of the ribozyme to be formed is the BP Helix: SF3A1 displaces SF1 and recruits U2 snRNP, U2 snRNA searches for homology by branchpoint-interacting stem-loop (BSL) which establishes a ‘toehold’ immediately upstream of the BP A and the BP Helix is formed in place of the BSL stem-loop by RNA strand invasion mechanism (Figure 6C). Based on reviewed structural, biochemical, mutation and phylogenetic evidence (Sickmier et al., 2006; Biancon et al., 2022; Yoshida et al., 2020; Kent et al., 2003; Irimia and Roy, 2008; Corrionero et al., 2011) we suggest, that U2 snRNA bridges across the PPT and fixes BP Helix by the conserved G=C pair with intron position -3. SF3B1 of the large U2-associated SF3A/B protein complex supports U2 snRNA base-pairing with variant BP sequences. Transition to complex A requires mutual stabilisation of U1 and U2 snRNPs: 5’ss and 3’ss are matched which is mediated by U2-associated SF3A1 binding to U1 stem-loop SL4, and UAP56 or SRSF1 that bridge U1 SL3 and U2AF65 or 35 respectively. At this stage cryptic sites are rejected, and the choice of alternative sites can be influenced by RBPs. By contrast, minor introns feature well-conserved splice sites, that do not admit for alternatives and in the minor spliceosome assembly U11 and U12 come together as di-snRNP, skipping the initial pre-selection and ensuring stable binding to both 5’ss and 3’ss at once. Precatalytic spliceosomes: U5•U6/U4 tri-snRNP joins the Pre-B complex and U2 5’ end pairs with U6 forming U6/U2 Helix II (Figure 6B). U6 ACAGAGA sequence is looped out between U6/U4 stem III and U4 quasi-pseudoknot supported by RBM42 bound to U4 opposite the loop (Figure 6A). Transition to B complex involves U1 snRNP disassociation from the 5’ss. Prp28 does not unwind the RNA duplex but instead destabilises it by interfering with the supporting protein U1C (Chen et al., 2001). U6 ACAGAGA loop binds the start of the intron, which leads to the disruption of the quasi-pseudoknot, removal of RBM42 and unwinding of U4/U6 stem III. Brr2 helicase loads to its binding site on U4 in place of stem III. All structural studies agree that at this stage 5’exon binds U5 Loop 1. Here we point out protosplice site repeats CAG|GU, conserved at both 5’ss and 3’ss. A robust mechanism is required to choose the correct junction of exons within repeats. The intron ends G+1••G-1 pair provides an essential link that allows the ribozyme to reposition intermediates between the two steps of splicing and to function in this way, it needs to be formed at pre-catalytic stage. Here we argue that the early involvement of the 5’ss-3’ss pair also provides a solution to the problem of operating with repeats. Splicesome structures do not include 3’ss at the pre-catalytic stage, here we follow the pre-catalytic structure of the Ll.LtrB Group IIA intron, that shows both 5’ss and 3’ss in proximity and the 3’exon aligned with the recognition loop (Liu et al., 2020). We showed previously (Artemyeva-Isman and Porter, 2021) that human U5 and U6 mutually compensate for variations in their binding sites: extensive U5 base-pairing with the 5’exon can support U6 binding to divergent introns and vice versa. This mechanism was cited and confirmed by later studies of different eukaryotes (Ishigami et al., 2021; Parker et al., 2022). We suggested the same mechanism for U5 and U2 at the 3’ss. The importance of specific U5 base-pairing with the 3’exon explained the effect of human mutations on splicing (Artemyeva-Isman and Porter, 2021) and transpired in later phylogenetic and mutation studies (Olthof et al., 2024; Negi et al., 2025; Nava et al., 2025). It is clear, that not only the U6 pairing with the intron, but overall stability of U6, U5 and U2 base-pairing with the pre-mRNA is the fidelity check at the pre-catalytic stage. Transition to the catalytic spliceosome starts with Brr2 helicase translocating on U4 in the 3’ to 5’ direction unwinding U4/U6 stems I and II. The catalytic triple helix of U2/U6 and U6 ISL are formed to coordinate the two Mg2+ ions. NTC/NTR (Nineteen complex, Prp19-associated proteins) join the complex to support the catalytic site conformation (Bact complex) and finally SF3A/B complex disassociates: SF3B1 quits the BP Helix. The bulged adenosine is clumped from Watson-Crick and Hoogsteen edges by the conserved U in position BP-2 and the invariant 3’ss A-2, respectively (Figure 2A). In this configuration the sugar edge of the BP A is facing outwards and ready for 2’O to attack 5’ss (complex B*, here designated C Step 1). Branching triggers a rotation of the BP Helix with the lariat intermediate, the 5’ss-3’ss pair serves to tow the 3’ss and the 3’ exon into the catalytic core. This conformational change is eased by the removal of the 1st step splicing factors. At the 2nd step 5’exon 3’OH attacks the 3’ss, the exons ligate and the intron lariat is freed, C Step 2 complex has the addition of the 2nd lot of splicing factors for conformational support. Next complex P (post-catalytic) still has the exons paired with U5 Loop1, but Prp22 facilitates mRNA release and the complex transitions to intron-lariat spliceosome ILS, that still has the catalytic site, U2 and U6 paired with the intron lariat and U5 snRNA with Loop1 unpaired and in heptaloop closed conformation. ILS is homologous to mobile Group IIA intron particle which can invade new loci by reverse splicing. SF1- Splicing Factor 1; U2AF- U2 Auxiliary Factor 65/35 heterodimer; Prp proteins – Pre-mRNA Processing; Brr2helicase (Bad Response to Refrigeration in yeast mutants) is U5-associated, binding site is indicated on U4 snRNA; 1st Step splicing factors: Cwc25, Yju2, Isy1; 2nd Step factors: FAM32A, Cactin, NKAP and Slu7 (Wilkinson et al., 2017).
Preprints 233649 g001
Figure 6. Intron splice site recognition by snRNAs: formation of the 5’ss and BP Helices. Splice sites and the BP site are highlighted in yellow. Figure panels A to C are arranged 5’ to 3’ on the pre-mRNA strand but need to be viewed C to A in order of ribozyme folding. See Figure 5 for successive spliceosomal complexes (E to ILS). A. 5’ss is initially selected by U1 snRNA with U1C protein acting as a supporting chaperon. In the pre-catalytic spliceosome (complex Pre-B) U5•U6/U4 tri-snRNP arrives, U6 is extensively paired to U4 (stems I, II and III) with a U4 turn (Quasi-pseudoknot) and RBM42 protein configuring ACAGAG loop (Charenton et al., 2019). Prp28 alters U1C conformation destabilising U1 base-pairing (Chen et al., 2001) which allows U6 ACAGAG loop to displace U1 from the start of the intron. 5’ss Helix extends unwinding U4/U6 Stem III, displacing RBM42 and disrupting the pseudoknot. Variations at splice sites form mismatched pairs; 5’ss Helix is stabilised by stacking of U6 Am643. Notably, m6 modification is not needed in homologous U6atac, that processes invariant minor introns (0.4% in humans, see Part 1, Artemyeva-Isman, 2026a - Figure 2A). Brr2 helicase loads onto U4 prepared to unwind stems I and II, if the pre-catalytic complex B passes the fidelity check: the attainment of overall stability in U6, U5 and U2 base-pairing with the pre-mRNA. U6/U2 Helix III – see below. B. U6/U2 Helix II is the first base-pairing interaction in the pre-catalytic spliceosome: it forms in Pre-B complex before 5’ss Helix. U6 snRNA never acts without U2. C. 3’ss, PPT and branchpoint BP site are initially selected together by U2AF35/65 and SF1. U2AF65 bends the PPT (Sickmier et al., 2006; Chen et al., 2010) juxtaposing BP and 3’ss (Kent et al., 2003). SF1 recruits and is displaced by SF3B1 component of U2 snRNP. U2 snRNA comes prepared with a branchpoint-interacting stem-loop (BSL) to find a ‘toehold’ at the most conserved UBP-2 (2nt 5’ from BP A). Intron strand invasion and BSL stem opening is followed by extension over the BP A (Cretu et al., 2021). -3C G mutation blocks complex A formation, but not U2AF binding (Corrionero et al., 2011), indicating C-3 involvement in the BP Helix, like in eukaryotes without a spacer between the BP site and 3’ss (Yarrowia lipolytica, Irimia and Roy, 2008). U2AF35 has no preference for C-3 (Yoshida et al., 2020), conservation of C-3 can be well explained by U2 snRNA bridging across the spacer forming a G=C pair, that brings BP at a fixed distance of 6nt from the 3’ss. BP Helix continuity to 3’ss is also a feature of homologous domain VI (DVI) in prokaryotic Group II introns, the ancestors of spliceosomal introns (Figure 2B in Part 1, Artemyeva-Isman, 2026a). Like 5’ss Helix, BP Helix is also capped with U2 Am6,m30 to stabilise mismatched pairs and homologous U12 does not need this modification for well conserved minor introns (Figure 2A). * U6/U2 Helix III was predicted between U6 nt 30-37 and U2 nt 42-49, but its existence is questioned, as it was not confirmed experimentally except for possibly U6 nt 30,31 and U2 nt 48,49 (see Part 4, Artemyeva-Isman, 2026d). Its existence and length clearly depend on the extent of 5’ss and BP Helices. CryoEM studies modelled both 5’ss and BP Helices extending further than a few conserved intron residues: 5’ss Helix until U6 A30 and BP Helix until U2 Ψ44 (Bertram et al., 2017a; Cretu et al., 2021). Early mutation studies affirm increase of splicing efficiency with 5’ss Helix extending to U6 position 38 (intron position +9, Crispino and Sharp, 1995; Hwang and Cohen, 1996). It is possible that the extend of 5’ss and BP Helices depends on the intron sequence and unpaired U6 and U2 bases not included in 5’ss and BP Helices will form a Helix III of variable length for additional support of the precatalytic complex.
Figure 6. Intron splice site recognition by snRNAs: formation of the 5’ss and BP Helices. Splice sites and the BP site are highlighted in yellow. Figure panels A to C are arranged 5’ to 3’ on the pre-mRNA strand but need to be viewed C to A in order of ribozyme folding. See Figure 5 for successive spliceosomal complexes (E to ILS). A. 5’ss is initially selected by U1 snRNA with U1C protein acting as a supporting chaperon. In the pre-catalytic spliceosome (complex Pre-B) U5•U6/U4 tri-snRNP arrives, U6 is extensively paired to U4 (stems I, II and III) with a U4 turn (Quasi-pseudoknot) and RBM42 protein configuring ACAGAG loop (Charenton et al., 2019). Prp28 alters U1C conformation destabilising U1 base-pairing (Chen et al., 2001) which allows U6 ACAGAG loop to displace U1 from the start of the intron. 5’ss Helix extends unwinding U4/U6 Stem III, displacing RBM42 and disrupting the pseudoknot. Variations at splice sites form mismatched pairs; 5’ss Helix is stabilised by stacking of U6 Am643. Notably, m6 modification is not needed in homologous U6atac, that processes invariant minor introns (0.4% in humans, see Part 1, Artemyeva-Isman, 2026a - Figure 2A). Brr2 helicase loads onto U4 prepared to unwind stems I and II, if the pre-catalytic complex B passes the fidelity check: the attainment of overall stability in U6, U5 and U2 base-pairing with the pre-mRNA. U6/U2 Helix III – see below. B. U6/U2 Helix II is the first base-pairing interaction in the pre-catalytic spliceosome: it forms in Pre-B complex before 5’ss Helix. U6 snRNA never acts without U2. C. 3’ss, PPT and branchpoint BP site are initially selected together by U2AF35/65 and SF1. U2AF65 bends the PPT (Sickmier et al., 2006; Chen et al., 2010) juxtaposing BP and 3’ss (Kent et al., 2003). SF1 recruits and is displaced by SF3B1 component of U2 snRNP. U2 snRNA comes prepared with a branchpoint-interacting stem-loop (BSL) to find a ‘toehold’ at the most conserved UBP-2 (2nt 5’ from BP A). Intron strand invasion and BSL stem opening is followed by extension over the BP A (Cretu et al., 2021). -3C G mutation blocks complex A formation, but not U2AF binding (Corrionero et al., 2011), indicating C-3 involvement in the BP Helix, like in eukaryotes without a spacer between the BP site and 3’ss (Yarrowia lipolytica, Irimia and Roy, 2008). U2AF35 has no preference for C-3 (Yoshida et al., 2020), conservation of C-3 can be well explained by U2 snRNA bridging across the spacer forming a G=C pair, that brings BP at a fixed distance of 6nt from the 3’ss. BP Helix continuity to 3’ss is also a feature of homologous domain VI (DVI) in prokaryotic Group II introns, the ancestors of spliceosomal introns (Figure 2B in Part 1, Artemyeva-Isman, 2026a). Like 5’ss Helix, BP Helix is also capped with U2 Am6,m30 to stabilise mismatched pairs and homologous U12 does not need this modification for well conserved minor introns (Figure 2A). * U6/U2 Helix III was predicted between U6 nt 30-37 and U2 nt 42-49, but its existence is questioned, as it was not confirmed experimentally except for possibly U6 nt 30,31 and U2 nt 48,49 (see Part 4, Artemyeva-Isman, 2026d). Its existence and length clearly depend on the extent of 5’ss and BP Helices. CryoEM studies modelled both 5’ss and BP Helices extending further than a few conserved intron residues: 5’ss Helix until U6 A30 and BP Helix until U2 Ψ44 (Bertram et al., 2017a; Cretu et al., 2021). Early mutation studies affirm increase of splicing efficiency with 5’ss Helix extending to U6 position 38 (intron position +9, Crispino and Sharp, 1995; Hwang and Cohen, 1996). It is possible that the extend of 5’ss and BP Helices depends on the intron sequence and unpaired U6 and U2 bases not included in 5’ss and BP Helices will form a Helix III of variable length for additional support of the precatalytic complex.
Preprints 233649 g002

2.2. 3’ss Selection by Protein Co-Factors SF1 and U2AF65/35

At the 3’ss the uncertainty of every splicing event is increased by the highly variable position of the branchpoint (BP): in human introns the estimated average is -25 (Paggi and Bejerano, 2018; Mercer et al., 2015; Briese et al., 2019), but it can be as far as -400 (experimentally tested distant branchpoints in Gooding et al., 2006). The length of RNA separating BP from the 3’ss usually contains pyrimidine stretches termed polypyrimidine tract (PPT). The sequence flanking BP adenosine is also highly variable in humans, but at the end of intron A-2 is invariant, and G-1 is 99% conserved (remaining 1% is C-1 in AT...AC introns). Therefore, 3’ss, PPT and BP are initially recognised cooperatively by three interacting proteins – based on the combination of an AG dinucleotide, a possible BP A and a PPT in the spacer between them. The mammalian splicing factor 1/ Branchpoint binding protein (SF1/BBP) binds BP A and the flanking sequence. The other two proteins are the subunits of the U2 Auxiliary Factor heterodimer: the larger (subunit 2) U2AF65 binds PPT and the smaller (subunit 1) U2AF35 binds the intron end AG.
U2AF65 crystal structures with polyU and variant strands show that uridines are preferred over cytosines, and an inclusion of a purine can be well tolerated (Glasser et al., 2022). The most interesting fact is that U2AF65 sharply bends the RNA strand with a kink of 1200 (Sickmier et al., 2006). Moreover, substitution of uridine for pseudouridine (Ψ) blocks U2AF65 binding, as the RNA strand is unable to bend (Chen et al., 2010). Its rigidity is conveyed by a water bridge: N1-H of Ψ coordinates a water molecule that forms hydrogen bonds with the 5’ phosphates of both Ψ and the preceding residue (Charette and Gray, 2000). U2AF65 bending PPT can explain how this exclusively protein-binding site is looped out of the tight ribozyme core.
U2AF35 affinity to A-2G-1 depends on the flanking nucleotides in intron position -3 and exon position +1. In fact, it is the only protein splicing factor for which specific amino acids are identified as having nucleotide preferences at specific positions with absolute certainty (Biancon et al., 2022). In myelodysplastic syndromes (MDS), frequent U2AF35 S34F mutation disrupts the splicing of exons preceded by 3’ss that lack C-3 conserved at 65% in human introns. Accordingly, S34F splicing pattern shows suppression of exons downstream of U-3, because it accounts for another 30% of human introns. The 5% left for purines will not show up on the logo (Figure 2C in Part 1 - Artemyeva-Isman, 2026a). Surprisingly, a recent X-ray crystallography and NMR study shows that WT S34 prefers the introns ending with UAG to CAG (Yoshida et al., 2020). We can conclude that C-3 is not conserved because of U2AF35 binding, but for some other later interaction, and that U2AF35 prompts or supports this other interaction for suboptimal U-3 and possibly rare purines as well. Similarly, the recurrent MDS Q157R mutation has an aberrant splicing pattern for exons that lack 50% conserved G+1 (Biancon et al., 2022), indicating that WT Q157 supports interactions with suboptimal nucleotides in the first position of exons. Therefore, U2AF35 helps to recognise variant 3’ss with unusual residues at intron position -3 and exon position +1.
SF1 binding (Peled-Zehavi et al., 2001 - structural model) of the vaguely conserved BP sequence clearly relies on U2AF65/35 recognition of PPT and 3’ss. The distance from the 3’ss and the nucleotide composition of the spacer with the PPT are possibly the main criteria for the initial selection.
3’ss and the adenine base intended for branching draw nearer in 3D with U2AF65 acting like a cable clip. Kent et al., 2003 proposed a model based on Fe-EDTA probing of the U2AF65-RNA interactions and a comparison with X-ray structures of related RNA Recognition Motifs (RRM): “U2AF65 bends the RNA to juxtapose the branch and 3′splice site”.

2.3. Branchpoint Pairing with U2 snRNA

The BP/U2 helix made in trans by two separate RNA molecules in the spliceosome is homologous to Domain VI stem of the Group II intron molecule (Figure 2A,B in Part 1 - Artemyeva-Isman, 2026a). In the spliceosome BP helix is the first structure of the ribozyme to form. Initially (complex E, Figure 5) BP site is bound by SF1, which in turn is displaced by SF3A1 (Nameki et al., 2022), a protein of subcomplex A of the large SF3A/B complex associated with U2. Transitioning to the pre-spliceosome (complex A in Figure 5) involves the U2 snRNA pairing with a sequence flanking BP in the intron. A brilliant mutation study in S. cereveisiae (Perriman and Ares, 2010) identified the branchpoint-interacting stem loop (BSL), which appears invariant in the U2 snRNA of all eukaryotes (Figure 6C). In humans the U2 28C U mutation, which destabilises BSL is frequently linked to haemotological, prostate and pancreatic cancers (Bousquets-Muñoz et al., 2022). According to CryoEM (Cretu et al., 2021), U2 initially conducts a homology search by the last trinucleotide of the BSL pentaloop G33Ψ34A35G36Ψ37, which establishes a ‘toehold’ at the BP site. Subsequently, the BSL stem unwinds and U2/BP helix extends 5’ to 3’ along the U2 strand, as this is the direction favoured by the thermodynamics of the toehold-mediated RNA strand displacement (Šulc et al., 2015). Following the formation of a stable helix upstream of the BP, it is completed downstream, bridging the bulged BP A, which is sterically recognised by SF3B subcomplex (Zhang Y. et al., 2025). Its largest subunit SF3B1 is frequently mutated in myelodysplasia and other cancers with widespread aberrant BP selection underlining a few critical alterations in 3’ss usage (Damianov et al., 2025). The ‘silent’ changes involved the use of a different BP linked to the same 3’ss. Remarkably, both with ‘silent’ changes and especially with altered BP usage linked to aberrant 3’ss, SF3B1 K700E mutant cells display improved BP site consensus. Significantly more Watson-Crick pairs occurred at positions BP-2 and BP-3, corresponding to the initial U2 ‘toehold’. This data indicates that WT SF3B1 supports suboptimal base-pairing of U2 with variable human BP sites.
Human U2 snRNA coordinates the bulged A between pseudouridine Ψ34 and guanine G33. Remarkably, the S. cerevisiae U2 and human U12, both of which have a lot fewer base modifications, share the U Ψ conversions in homologous positions supporting the bulged adenosine. Apart from conveying rigidity to the sugar-phosphate chain 5’ from pseudouridylation site (as explained in the previous section) Ψ-A pair is extra stable, and at the same time Ψ like U can form pairs isosteric to Watson-Crick with any other base to compensate for variations.
The extent of the BP helix in human introns is impossible to define by a clear consensus, as around BP A nt conservation is reduced to 30-37%, except for 60% for UBP-2 for a later stage interaction optimising catalysis (see 4.1 BP helix preparation for branching). We can infer human BP helix from the well conserved introns of S. cerevisiae, because the U2 site, pairing with the intron is invariant in all eukaryotes (see Part 1, Artemyeva-Isman, 2026a). U2 U32G33_Ψ34A35G36Ψ37A38 (human nt numbering) matches the S.c. BP motif: UBP-5 ACUAACABP+2 (where A is the BP A).
ABP+2 is often excluded from the BP motif, which is inaccurate despite its poor conservation. It aligns with human U2 U32; uracil in homologous U2 positions is conserved throughout eukaryotic taxa. A rare exception, C. merolae U2 has a U G change in a homologous position (Stark et al., 2015). Accordingly, C.m. introns have a reciprocal change for CBP+2 (Slat, 2021), making a G=C pair (Stark et al., 2015). This exceptional case of a different Watson-Crick pair produced by co-variations in U2 and BP sites serves to confirm the architype U-A pair (human U2 U32-ABP+2) despite variations at BP sites of human major introns. In human minor (U12) introns this pair is swapped between the interacting RNAs: conserved UBP+2 (Olthof et al., 2024) pairs with U12 A17 (Figure 2A, Minor, U12 intron – BP helix).
As we already remarked above, in human introns, the spacer between BP and 3’ss is extremely variable in length and pyrimidine/purine composition: on average there are some 20nt between ABP+2 and C-3 (Figure 2A in Part 1 - Artemyeva-Isman, 2026a). This is in sharp contrast to Group II introns where the distance between BP A and 3’ss is fixed to 6-7nt (Costa, 2022) and the BP helix extends without interruption to intron C-3 (Figure 2B). Notable exceptions were found in Bacillus sp.: Group IIB introns with insertions of 53-70nt separating BP from 3’ss (Tourasse et al., 2011). This stretch is predicted to fold into two stem loops, an additional domain VII. While in spliceosomal introns PPT is looped out by U2AF65 (see previous section), these Group IIB introns rely on RNA base-pairing to ensure that the long spacer does not disrupt the ribozyme core. Exceptionally, there are unicellular eukaryotes that do not have a spacer between BP and 3’ss (Xie et al., 2023). The smallest fixed BP to 3’ss distance of 6nt is found in yeast Yarrowia lipolytica (Irimia and Roy, 2008). The well-conserved Y. lipolytica BP motif appears adjacent to the 3’ss, so its introns end with a consensus UACUAA-6CAC-3AG|, where A-6 is the BP A. The singular fact is that in Y. lipolytica the BP motif and C-3 of the 3’ss can bind continuously with U2 G31U32G33_Ψ34A35G36Ψ37A38 (human nt numbering – Y.l. U2 has an identical sequence as do most eukaryotes). It appears that the minimum distance between BP and 3’ss is simply defined by the U2 binding register, and that continuous BP helix to 3’ss can occur in the spliceosome like in Group II introns (Figure 2B). Indeed, a spacer between BP and the 3’ss is so disruptive to ribozyme assembly, that we might want to inquire further into U2 snRNA base-pairing with the intron. The strong conservation of the human intron position -3: 65% C-3 and 30% U-3 is consistent with almost exclusive variation of G=C/G--U pairs (possibly a mimic pair, isosteric to Watson-Crick, rather than G•U wobble, which has a different shape). It cannot be ignored that U2AF35 binding does not account for C-3 conservation and preferably supports U-3 (see previous section). The solution for keeping the ribozyme core together can be for U2 snRNA to bridge over the interrupting spacer and base-pair with C-3 bringing the BP helix to the 3’ss as in Group II introns. The ends of the spacer are closer in 3D, as U2AF65 interacting with PPT sharply bends the RNA strand (see previous section). Human U2 G31 (U12 G16) is there as a candidate for intron C-3 partner (Figure 2A). This guanine is universally conserved in homologous U2 (U12) positions in eukaryotes. An excellent mutation study on FAS intron 5 -3C G in Autoimmune Lymphoproliferative Syndrome ALPS (Corrionero et al., 2011) showed that U2AF initial binding was not affected by this mutation, but the formation of complex A was blocked, indicating -3C involvement in the BP helix. Another relevant observation comes from analyses of frequent 3’ splice sites within NAGNAG motifs. These are found in some 30% of human genes and ~5% of these are involved in functional alternative splicing, estimated as 1/5 of all alternative splicing events. Bradley et al., 2012 convincingly show that ‘the -3 bases largely determine whether a NAGNAG is alternatively spliced’, again highlighting the importance of C-3/U-3 conservation, for which no mechanistic explanation has been put forward so far. The genetic and structural evidence reviewed here suggests that juxtaposition of the BP site with 3’ss, achieved by U2AF and SF1 at the previous stage (Kent et al., 2003, see previous section) is fixed by the inferred U2 G31=C-3/G31--U-3 pair. As U2 is bridging over the variable spacer, BP A is fixed 6nt from the 3’ss, like in the continuous BP-3’ss helix in yeast Y. lipolytica and the BP helix (DVI stem) in Group II introns (Figure 2A,B, experiments to test this hypothesis are outlined in Part 4 - Artemyeva-Isman, 2026d).

2.4. 5’ and 3’ss Coupling and RBP Activity

Transition to pre-spliceosome (complex A, Figure 5) depends on mutual stabilisation of U1 and U2 snRNA at splice sites. U2/BP-3’ss gets matched to a suitable U1/5’ss, while the surplus of cryptic 5’ss and ectopic 3’ss/PPT/BP-like sequences are rejected depending on their affinity to U1 and U2 respectively, and their proximity to each other. Stabilisation interactions between them appear to be mediated by U1 stem-loop 4 SL4, which binds the U2 snRNA associated protein SF3A1, homologue of yeast Prp21, while UAP56/DDX39B helicase bridges between U1 SL3 and U2AF65 (Martelly et al., 2021). In addition, SRSF1, bound to exonic splicing enhancers ESE, is also known to interact with both U1 SL3 and U2 snRNP, possibly via U2AF35 (Jobbins et al., 2022). Alternative combinations of 5’ and 3’ss are formed in proportion to their stability, but the frequency of these choices can be skewed to favour a particular mRNA isoform by regulatory RNA binding proteins RBP. Expansion of SR and hnRNP in animals and plants was already mentioned (see Part 1, Artemyeva-Isman, 2026a), but these are only two families of ubiquitous splicing factors (SF). Many more tissue-specific SF are required for cell differentiation and stress response.
Studies linking SF with diseases overflow current literature and it is increasingly apparent that the mechanisms of RBP-mediated splicing regulation are diverse. Generally, SR proteins bound to exonic enhancers (ESE) bridge between U2 and U1 snRNPs and stabilise early complexes across exons promoting splicing, while hnRNP proteins bind intronic silencers (ISS) and block snRNP binding or prevent interactions between snRNPs suppressing splicing. However, this explanation is oversimplistic, because the action of RBPs is governed by positional effects. Examples show that depending on binding location(s) in the adjacent introns ubiquitous PTBP1 (hnRNPI, Hamid and Makeyev, 2017) and neuronal-specific NOVA (Ule and Belncowe, 2019) can enhance or suppress exon inclusion. For PTBP1 this can be simply a question of expression level, which affects the occupation of sites with less affinity. Furthermore, RBPs with antagonistic effects on exon inclusion can compete to produce precise tissue-specific ratio of splicing isoforms. Alternative splicing to the maximum complexity in case of titin TTN, a transcript of 364 exons, produces distinct groups of skeletal muscle and cardiac isoforms, that exclude many exons that are used only in embryonic development (Savarese et al., 2016, Gregorich et al., 2024). The two main SF of TTN are the RNA-binding motif protein 20 RBM20 and Quaking QKI (Montanes-Agudo et al., 2023); both generally act as suppressors of exon inclusion. QKI is recently discovered to block BP sites competing with SF1 (Pereira de Castro et al., 2024), and Qki-/- is embryonic lethal in mice (Sakers et al., 2021). RBM20 mutations cause inclusion of embryonic and skeletal muscle-specific TTN exons in the heart, and, in addition, splicing dysregulation of a few other cardiac transcripts, leading to severe dilated cardiomyopathy DCM and heart failure (Gregorich et al., 2024). Furthermore, expression of RBM20 is controlled by circadian factors BMAL1 and CLOCK and their function is altered in aging (Riley et al., 2022).
Despite their apparent regulatory importance RBPs recognise only 3-6nt of degenerate sequence. The best proposed explanation involves ‘multivalent’ RNA-protein interactions: multiple RNA sites are bound by RNA Recognition Motifs (RRMs) of the same protein – PTBP1 has 4 RRMs, NOVA has 3 RRMs – or by several protein molecules with single RRMs– like hnRNPC homotetramers or larger assemblies of different RBPs connected by their Intrinsic Disordered Regions IDR. Protein-RNA interactions are further stabilised by molecular crowding, as multivalent interactions promote liquid-lipid phase separation LLPS (Ule and Blencowe, 2019). This characteristic RBP-RNA accretion is pathogenic on mutant transcripts with expanded repeats in neurodegenerative diseases (Depienne and Mandel, 2021). For example, in myotonic dystrophy SF MBNL1 and 2 are sequestered on DMPK CUG or CNBP CCUG expanded repeats and accumulate in LLPS nuclear foci (Sellier et al., 2018). Aberrant splicing of several transcripts critical at neuromuscular junctions is caused by the lack of free MBNL proteins (Frison-Roche et al., 2025).
It is important to see the principal difference between RNA base pairing and RBP binding in splice site selection. The first is the start of ribozyme assembly guided by clear rules of nucleic acids complementarity; the second is of limited specificity and helps or prevents RNA assembly. It is therefore not surprising that RBP regulation is relevant to ‘weak’ splice sites only (Hamid and Makeyev, 2017; Fukumura et al., 2009) and mutant splice sites with higher affinity to snRNAs are chosen regardless of RBPs. The same can be achieved if snRNAs are engineered to increase affinity to specific splice sites (see Parts 3 and 4 Artemyeva-Isman, 2026c,d).
The minor spliceosome, which is likely closer to the ancestral form (see Part 1, Artemyeva-Isman, 2026a), has U11 and U12 snRNPs joined together by protein interactions forming di-snRNP to ensure coordinated 5’ and 3’ss selection. This fixed arrangement and well-conserved U11 and U12 binding sites explains little alternative splicing of minor introns. However, as already mentioned, some AS of minor introns can be explained by assembly of hybrid spliceosomes (Olthof et al., 2024). In such cases the early complexes must be skipped. As reflected previously, minor or major introns are distinguished by their partners U11 or U1 and U6atac or U6 snRNAs on the basis of a G=C pair swapped between the interacting RNAs at intron position +5 (Figure 2A). As U11 and U12 act together in the early complexes, the choice of the BP site is linked to U11 base-pairing, and no base swapping between the interacting RNAs are required in the BP helix. Remarkably, conserved 5’ss and BP sites in minor introns have doubled G=C pairs (Figure 2A) – a feature that disappears in major introns that drifted towards minimal required base-pairing to allow for alternative splice site selection in the early complexes (E and A).

3. Pre-Catalytic Spliceosome: Fidelity Check

In the pre-catalytic spliceosome (Figure 5) U5•U6/U4 tri-snRNP joins U2 snRNP (complex pre-B). To prevent premature formation of the catalytic site, the U6 snRNA comes extensively paired to U4 snRNA (Figure 6A). The free 3’ end of U6 pairs the 5’ end of U2 forming U2/U6 helix II – a strong 9bp duplex of 6 C=G/G=C pairs, 2 U-A and 1 U--U mismatch (Figure 6B). Transition to complex B releases U1 snRNP from the 5’ss, which is facilitated by Prp28. Instead of using its helicase activity, Prp28 interferes with the U1C role as a chaperone of the U1/5’ss helix (Chen et al., 2001), leading to destabilisation of the RNA duplex, so U1 is replaced by the competing U6 to form the 5’ss helix. For this purpose, the U6 ACAGAGA sequence is looped out between U4/U6 stem III and U4 quasi-pseudoknot, chaperoned by RBM42 (Charenton et al., 2019). The quasi-pseudoknot is a fold involving U4 nt63-67 secured by two non-canonical base-pairs: U4 U63••A67 in tWH configuration and U6 A47 (the last A of the ACAGAGA box) with U4 G64 as tWS pair (Figure 6A; Charenton et al., 2019 - structure in supplementary material, see Leontis et al., 2002 for Westhof geometric classification of RNA base pairs1). U4 mutations in the region of the pseudoknot, RBM42 binding site and U4/U6 stem III cause ReNU neurodevelopmental syndrome (Chen et al., 2024; Greene et al., 2024; Nava et al., 2025); a plausible mechanism is an altered U6 ACAGAGA loop conformation that is deficient in recognition of introns, favouring 5’ss flanked by better conserved exons (Nava et al., 2025 – see below 3.3 Exon recognition by U5 snRNA Loop1). U4/U6 stem III is generally conserved in eukaryotes and apart from demarcating the U6 ACAGAGA loop, plays an important role preventing Brr2 helicase from loading onto U4 snRNA at the pre-B stage. The homologous U4atac/U6atac stem does the same. Although stem III was first discovered when predicting the secondary structure for U4/U6 of a monocellular eukaryote (green algae Chlamydomonas reinhardtii – Jakab et al., 1997), it appears to be absent in yeast S. cerevisiae U4/U6 – one of many indications of the unfortunate choice of this organism as a model for studying the mechanism of splicing. S.c. Brr2 comes loaded on U4 snRNA in the pre-B complex. However, in the usual assembly pathway, including humans, U6 ACAGAGA loop interaction with the 5’ss triggers U4/U6 stem III unwinding, which then frees Brr2 loading site on U4 snRNA. In fact, every precaution is taken to prevent premature splicing catalysis: pre-catalytic complexes evolved to safeguard splicing precision. Vaguely selected splice sites from complex A are now double checked and the formation of the catalytic site of the ribozyme is delayed by U4 snRNA while the check is completed, and only then Brr2 unties U6 from U4.

3.1. 5’ss Helix

In Group IIA introns (Figure 2B in Part 1 - Artemyeva-Isman, 2026a) 5’ss helix is designated as ε-ε’ interaction between the start of the intron and an internal loop of the Ic1 stem (part of Domain I). Ll.LtrB intron ribozyme has two certain pairs next to 5’ss, G+3=C109, C+4=G108, but two more consequent Watson-Crick pairs can be inferred between the 5’ss G+1U+2 and A110C111, which are the last two unpaired residues of the Ic1 internal loop. For a wider picture, in Group IIA introns A is almost invariant in homologous position to Ll.LtrB A110 and U is conserved instead of Ll.LtrB C111, suggesting a mimic G--U or a wobble G•U pair. By contrast, Ic1 loop of Group IIB or IIC introns do not include more than 2 or 1nt complementary to the start of the intron, limiting ε-ε’ (5’ helix, Figure 2B) to 2 or 1bp. The evolutionary distances between Group IIA and subclasses IIB, IIC are large enough to account for distinct mechanisms of 3’exon recognition and BP helix tertiary interactions, so it is not out of order to suggest that IIA 5’ss helix (ε-ε’) includes G+1U+2, despite it not being the case in IIB and IIC subclasses.
In Group II introns regardless of subclass, Watson-Crick pairs of the ε-ε’ element are immediately adjacent to a tertiary interaction λ-λ’ (Boudvillain et al., 2000) that in Ll.LtrB (Figure 2B) involves A107 and the minor groove of the C2407=G2418 of Domain V, tethering 5’ss to the catalytic site of the ribozyme. This type of interaction is termed A-minor, a very common motif in tertiary structures of ribozymes.
In spliceosomal introns, the 5’ss helix between the start of the intron and U6 snRNA has more in common with Group II introns than meets the eye at first glance (Figure 2A,B). An obvious difference is that the Watson-Crick pairs, starting with the invariant C+5 in minor introns or the conserved G+5 in major introns (78%, Figure 2C) are separated from the 5’ss by one or two adenines, which cannot form Watson-Crick pairs with U6: in minor introns invariant A+3 aligns with U6atac G17 (Figure 2A - minor intron 5’ss helix), while in major introns conserved A+3A+4 (60 and 71%, Figure 2C) align with U6 G44Am643 (Figure 2A). As we just reviewed, the Watson-Crick pairs ε-ε’ in the Group IIA Ll.LtrB intron are adjacent to (and probably include) 5’ss; G+5 aligned with A107 (λ) comes immediately after ε-ε’ (Figure 2B). This looks like an inversion in the 5’ss helix downstream from U+2. Although it is impossible to tell when and how this inversion occurred, it turned out an adaptive advantage for ribozyme assembly in trans and accommodation of poorly conserved splice sites, necessary for alternative splicing. In Group II introns 5’ss helix is very short, only 1-4 Watson-Crick pairs (ε-ε’), which works as an intramolecular interaction. By contrast, in the spliceosome, 5’ss helix is an intermolecular joint: the minor introns positions +4 to +9 and U6atac form six almost invariant Watson-Crick pairs; the major introns positions +5 to +9 and U6 can form Watson-Crick pairs (Crispino and Sharp,1995; Hwang and Cohen, 1996), although only G+5 is conserved. CryoEM structure of the human pre-catalytic spliceosome features an especially long 5’ss helix extending to U6 A30 – the intron position +17 (Bertram et al., 2017b). Prolonged 5’ss helix is a strong joint for the assembly in trans, which can stabilise mismatches to allow for alternative splice site choices. The A-minor (λ-λ’) link with the catalytic site would have been removed too far away from the 5’ss if not for the inversion which fixes it at +3 or +4. Human major introns have the conserved adenine repeated A+3A+4 to make sure that at least one is there despite variability of splice sites, while S. cerevisiae introns and human minor introns possess the invariant A+3, and U+4 forms the next Watson-Crick pair. In C. merolae this Watson-Crick pair is swapped between the interacting RNAs: A+4 pairs with unprecedented U6 ACUGAGA (Black et al., 2023), a change in the ‘ACAGAGA box’ invariant throughout eukaryotic taxa (see below). Another exception is the S. cerevisiae RPL30 intron with C+3A+4. This intron was examined by comprehensive mutation analyses, which included mutating the opposite U6 strand positions (Konarska et al., 2006). The study shows that A+4 indeed compensates for the lack of A+3. The conclusion was made that changing the adenines to make Watson-Crick pairs with U6 blocks the second step. However, there was also a dramatic reduction of the first step efficiency by a factor of 100 in substrates lacking both adenines. We may conclude more broadly that at least one conserved adenine at the start of intron position +3 and/or +4 is necessary and sufficient for normal ribozyme function.
Conservation of the A-minor motif in the 5’ss helix of spliceosomal introns is further supported by the presence of a C=G pair in the homologous position of U6(atac) ISL as in Group IIA Domain V: U6 Cm63=G71, U6atac C35=G43, and Ll.LtrB intron C2407=G2418 (Figure 2A,B in Part 1 - Artemyeva-Isman, 2026a). For U6 Cm63 2’O-methylation of the ribose increases base-pairing stability by improving sugar pucker stacking of the stem. On the contrary, the U6 Cm62=Gm272 pair is an impossible recipient for A-minor interaction. Gm272 -methylation of guanine C2-NH-CH3 does not interfere with Watson-Crick base-pairing, but the methyl group will be blocking the minor groove side of this pair. This is a device to support the specific position for A-minor link. However, while the recipient pair is intact, the inversion makes the conserved adenine switch strands (compare Figure 2A and 2B). Is this a problem for A-minor interaction? The answer is: unlikely. Adenine inserts into the minor groove of C=G and G=C with almost equal efficiency. More specifically, A-minor motifs examined in Group I introns with transversed C=G/G=C recipient pairs have a very small energetic difference of 0.4 kkal mol-1 measured by folding essays and very small rms deviations of 0.85 Å in crystal structures (Battle and Doudna, 2002).
However, despite repeated conserved adenines A+3A+4, there are rare human introns that lack them both. How does the ribozyme fold in these exceptional cases? A-minor motifs are abundant in ribozymes and appear in different structural contexts, due to this motif’s remarkably varied configuration. Four distinct types were identified among 168 A-minor interactions in the large ribosomal subunit (Nissen et al, 2001). While adenines are preferred in all A-minor types, other nucleotides can substitute adenine in weaker versions of this interaction (Nissen et al, 2001).
U6 position 43, the 6-methyl-adenine Am643 deserves special attention. Neither U6atac (Figure 2A), nor S.c. U6 have this modification; in both of those cases, the start of the intron is almost invariant. Two recent studies show that the absence of Am6 methylation in homologous U6 positions of yeast Schizosaccharomyces pombe and a flowering plant Arabidopsis thaliana leads to altered usage of 5’ splice sites and increased dependence on better conserved exons (Ishigami et al., 2021; Parker et al., 2022 – see below in 3.3 Exon recognition by U5 snRNA Loop 1). 6-methyl modification of adenine has two effects: on the one hand, it blocks Watson-Crick pairing with uridine, but it also stabilises the duplex overall by stronger stacking than by unmodified adenine. The general stabilising effect must be especially important for 5’ss helix if the start of the intron is variable as in human major introns (Figure 2C).
Group II ribozyme assembly with IEP in trans and expansion of IEP support functions eventually led to deletion of RNA Domains II, III and IV. Crucially, this provided a new covalent link between the 5’ss helix and the catalytic site. Indeed, splicing studies since 1990s discovered that U6 snRNA ‘ACAGAGA box’ plays a central role in spliceosome assembly: A41C42Am643G44A45 aligns and pairs with the start of the intron and later, after U4 snRNA releases U6, G46A47 is incorporated into the triple helix of the catalytic site. As intron ribozyme combines unique splice sites and universal snRNAs, it requires stringent control of recognition fidelity and U6 ACAGAGA bridge provides a very useful link to control catalytic site set-up depending on the 5’ss helix stability.
Group II intron folding combines parts of the same RNA molecule: 5’ss and BP helices do not have the recognition role like in spliceosomal introns. Looking for some homology to the ACAGAGA box in Group IIA Ll.LtrB intron (Figure 2 A,B in Part 1 - Artemyeva-Isman, 2026a), we trace its last GA dinucleotide, as the metal binding site is primeval: G471A472 is located at the junction between Domains II and III - J2/3 (Keating et al., 2010). Next in U6 is A45, which aligns with U+2. This is mirrored by the Ll.LtrB intron A110 likewise aligned to U+2, which appears to continue ε-ε’ Watson-Crick duplex (see above, Figure 2 A,B). The importance of U+2 is conveyed by its remarkable conservation: the ends are almost invariant GU...AG in spliceosomal and GU...AY in Group II introns. Early crosslinking studies suggested that U6 A45 and intron U+2 are paired throughout the splicing process (precatalytic stalled by prp2-1 heat inactivation - Bact, S.c.; lariat intermediate – C, catalytic step 2, S.c. and human; and lariat product – post-catalytic P, human – Kim and Abelson 1996, Fabrizio and Abelson 1990, Sontheimer and Steitz 1993). CryoEM reconstructions were not consistent on this account: this pair was absent in structures from Lührman (Rauhut et al., 2016; Bertram et al., 2017a,b; Haselbach et al., 2018) and Shi Labs (Wan et al., 2016a,b; Yan et al., 2016; Wan et al., 2019; Yan et al., 2017; Bai et al., 2017; Wan et al., 2017; Zhan et al., 2018a,b; Zhang X. et al., 2018; Zhang X. et al., 2017), while it was present in S.c and human structures from Nagai Lab as a trans-Watson-Crick-Hoogsteen tWH pair at precatalytic, catalytic step 1 and post-catalytic stages, but it was disrupted temporarily at catalytic step 2 stage (Charenton et al., 2019; Fica et al., 2017; Galej et al., 2016; Fica et al., 2019). This pair finally appeared in the structures from Shi Lab in human post-catalytic and intron-lariat spliceosomes (Zhang X. et al., 2019), despite its previous absence in analogous S.c. complexes. In view of the remarkable conservation of U+2 alignment with A110 in Group IIA Ll.LtrB structure (Figure 2B), the evidence from spliceosome crosslinking, and CryoEM structures from Nagai Lab, U+2 really pairs with U6 A45. The exact configuration can be disputed due to lack of more evidence. In 1% of human introns U+2 is substituted for C+2: although tWH U∙∙A and C∙∙A are nearly isosteric, a mimic C--A isosteric to canonical Watson-Crick U-A is also a possibility (more on Watson-Crick-like pairs in 3.3 Exon recognition by U5 snRNA Loop1). As for being temporarily disengaged at catalytic step 2 – this depends on the neighbours of U+2. The downstream neighbour A+3 was discussed at length above as a conserved A-minor motif (homologous to λ-λ’ in Group II introns). This interaction must be formed with the catalytic site and is certainly not disengaged between the two steps of splicing (according to biochemical and CryoEM studies of the Group IIA Ll.LtrB intron Dong et al., 2018; Liu et al., 2020). The upstream neighbour G+1, the conserved first guanine of the intron is engaged in the pairing of the intron ends (as a reminder: the ends are GU...AG in spliceosomal and GU...AY in Group II introns). The timing for this interaction will be specially discussed (see next section).
Returning to the tracing of Group II intron homology to the ACAGAGA box, let us now look at the differences. While U6 A45 is adjacent to G46Am47, in Group IIA Ll.LtrB intron there are 360nt (most of Domain I and the whole of Domain II) separating A110 from the catalytic triple helix (Figure 2 A,B). Instead of the U6 ACAGAGA bridge between the 5’ss helix and the catalytic site, Group II introns possess an isolated Watson-Crick pair γ-γ’, G470=C2492 in Ll.LtrB intron, which links the catalytic site with the 3’ss (GU...AC in Ll.LtrB intron). It is tempting to speculate that the connection of the catalytic site to 5’ss or 3’ss is structurally synonymous between Group IIA and spliceosomal introns, because 5’ and 3’ss in all lariat introns are joined together by a base pair of the intron ends (GU...AY in Group II introns andGU...AG in spliceosomal introns).

3.2. Pairing of the Intron Ends, 5’ and 3’ss Demarcation

Group IIA and spliceosomal introns maintain a tertiary interaction between 5’ and 3’ss to reposition splicing intermediates (see Part 1 - Artemyeva-Isman, 2026a). The commonly conserved G+1 is involved in a non-Watson-Crick pair with the 3’ss, linking intron ends in parallel orientation, although the exact configuration depends on the connection with the catalytic site, which is either the γ-γ’ pair in Group II introns or the U6 ACAGAGA bridge in spliceosomal introns. In Group II introns G+1 pairs with a penultimate base of the intron, usually A-2, found in introns of subclasses IIA, IIB and IIC. Mutation analyses of a Group IIB intron (Chanfreau and Jaquier, 1993) showed that any substitutions of G+1 reduce the efficacy of both branching (first step) and exon ligation (second step); A-2 mutations for G-2 and U-2 inhibited the second step, but combining G+1 and C-2 yielded 90% efficacy at the first step and almost 70% at the second step. The exact configuration of the G+1••A-2 pair was finally captured in the crystal structure of a Group IIC chimeric intron (Costa et al., 2016): the sugar edge of G+1 bonds with Watson-Crick edge of A-2 with glycosidic bonds in trans orientation and parallel local RNA strands orientation (tWS, 6th Westhof geometric family). Convincingly, G••C pair in this configuration is nearly isosteric to G••A. Moreover, if we infer that this is also the conformation of the Group IIA Ll.LtrB G+1••A-2 pair, then the Watson-Crick edge of G+1 is available to pair with C111 in the 5’ss helix (Figure 2B in Part 1 - Artemyeva-Isman, 2026a).
In spliceosomal introns U6 ACAGAGA bridges 5’ss to the catalytic site, functionally replacing the γ-γ’ interaction which linked 3’ss to the catalytic site in Group II introns. This is an acceptable change if the tertiary link between 5’ss and 3’ss is preserved. The loss of γ-γ’ also leaves intron position -1 free for the interaction between the ends of the intron. Accordingly, human intron ends are linked by a non-Watson-Crick pair between the first and last nucleotides (Parker and Siliciano, 1993; Chanfreau et al.,1994; Deirdre et al., 1995). This is why human introns are defined by guanines at their ends, and they technically cease to be introns if one of these guanines is mutated. The only other combination is A at the start and C at the end, which account for 1% of major introns and 13% of minor introns. The exact configuration of this pair was established by systematic mutation analyses (Deirdre et al., 1995). The only substitutions that support splicing are A+1••C-1, I+1••I-1 and to some degree A+1••A-1. Inosine (I) is a guanine analogue lacking the exocyclic N2-amino group, which means it does not provide hydrogen bonds for this pair. The only possible configuration with nearly isosteric G••G, A••C and A••A is the 2nd Westhof geometric family: bases interact by Watson-Crick edges, glycosidic bonds in trans and RNA strands in parallel orientation (tWW, termed N1-carbonyl symmetric in the original study). Unfortunately, CryoEM studies (Wilkinson et al., 2017; Bai et al. 2017; Fica et al., 2019; Zhang X. et al., 2019), despite referring to the same mutagenesis study (Deirdre et al., 1995), featured this pair as cis-Watson-Crick/ Hoogsteen (cWH), which is awkward, as this requires -NH2 involvement from the last guanine of the intron and A••C and A••A do not exist in cWH configuration.
Here we come to the exciting part: let us seek to answer three questions, central for lariat ribozymes (Figure 2A,B): (Q1) Why do spliceosomal introns have a different 5’-3’ss pair? (Q2) Why do the exons bind the recognition loop asymmetrically? And – (Q3) How are the cleavage sites precisely defined? Starting from question (Q1), we first look at the primary sequence conservation. It is remarkable, that the end G-1 never naturally occurs in Group II introns. The last nucleotide involved in the γ-γ’ base pair (see previous section) is as likely to be C-1 as U-1 (GU...AY), with one known exception ending with A-1 (Michel et al., 1989). Enigmatically, in vitro splicing of a mutated Group IIB intron (Jacquier and Michel, 1990) was unaffected when wt γ-γ’ A-U pair was substituted for γ-γ’ C=G pair (GU...AU  GU...AG). Further, we will deliberately not look at Group IIB and IIC subgroups, as they differ in their way of 3’exon recognition. As for Group IIA introns, like Ll.LtrB (Figure 2B) Michel and Ferat, 1995 observed that interactions in the ribozyme “prevent the first nucleotide of the intron from being in helical continuity with the 5' exon”. This is ensured by the primary sequence: if 3’exon happens to start with a G, the conserved intron G+1 is changed to an U+1 (|UU...AY|G), so that the C in Id3 Loop, that is meant to pair with the start of the 3’exon will not pair intron G+1 by mistake.
Spliceosomal introns strikingly differ in these primary sequence preferences: the intron end G-1 is almost invariant, and equally the invariant G+1 is combined with the conserved G at the start of the exon. Surprisingly, they can operate with repeats at their splice sites (AG|GU...AG|G), and the correct junction of exons is identical to these repeats (AG|G): the ancestral protosplice site (Figure 2A, see Part 1 - Artemyeva-Isman, 2026a). Accordingly, the U5 snRNA Loop1, which serves to align exons for ligation, like Group IIA Id3 Loop, is complementary to the protosplice site, U5 C38|C39U40. How is intron-start G+1 prevented from pairing with U5 C38 to avoid helical continuity across the exon/intron boundary? This problem is solved by the tWW 5’-3’ss pair which uses intron G+1 and G-1 Watson-Crick edges excluding them from binding U5 Loop1 along with the exons. In this way the intron ends are precisely defined within the splice site repeats and the reconstructed junction of exons AG|G is aligned on U5 Loop1. Evidently, this tWW 5’-3’ss pair coevolved with the ancestral intron that used protosplice site.
Here we must revisit intron evolution: admittedly, there is some uncertainty in the order of events that led to the massive proliferation of introns in eukaryogenesis. Is a successful Group IIA progenitor that spread duplicating protosplice sites as described in Part 1 (Artemyeva-Isman, 2026a) structurally possible? A Group IIA intron must have a γ-γ’ interaction, if this intron had the exceptional GU...AG ends, γ-γ’ was C=G, which is not found in modern Group II introns, but works experimentally (see above, Jacquier and Michel, 1990). If this intron had G+1••A-2 intron ends pair in tSW configuration, then it relied on intron G+1 Watson-Crick edge involvement in the 5’ss helix (see above 3.1 5’ss helix, Figure 2B) to prevent it from binding Id3 loop in helical continuity with the 5’exon. Intron start G+1 and 3’ exon start G+1 never occur together in modern Group IIA introns to prevent any ambiguity that can lead to intron start binding Id3 loop (see above, Jacquier and Michel, 1990). However, we might allow for an intron progenitor to be more of a ‘risktaker’ in early eukaryogenesis. The process of intron fragmentation and the evolution of assembly in trans led to the deletion of RNA Domains II, III and IV; the U6 ACAGAGA box linked 5’ss to the catalytic site (see above 3.1 5’ss helix) compensating for the loss of γ-γ’ tertiary link (Figure 2A,B). The change for the intron termini G+1••G-1 pair was the most effective solution for splicing CAG|GU repeats. Notably, 5’ss helix with U6 ACAGAGA decidedly excludes G+1, unlike in the Group IIA intron set up (Figure 2A,B): any ambiguity of base-pairing is now avoided in the ribozyme core. Alternatively, the intron progenitor and the protosplice site originated in a multistep process in eukaryogenesis. The intron progenitor that massively proliferated by reconstructing protosplice sites was already a fragmented ribozyme that had the intron termini G+1••G-1 pair. However, this behaviour can be attributed only to the simplest proto-spliceosome, as Vosseberg et al., 2023 show that the spread of introns took place before the emergence of a complex spliceosome.
Returning to our questions: a clue for a different 5’-3’ss pair in spliceosomal ribozymes appears to be a structural solution for splicing within repeats (Q1). We will return to the intron ends pair, the asymmetry of exons binding the loop (Q2) and definition of the cleavage sites by the ribozyme (Q3) in the next section, examining the 3D structure (Figure 7).
CryoEM models of the pre-catalytic 5’ss in the spliceosome focus on the interactions with Prp8 protein. Group II introns intimately co-evolved with their protein co-factor IEP, which supports ribozyme folding by stabilising the correct RNA conformation. Prp8, the main spliceosomal protein associated with U5 snRNP directly descended from the ancestral IEP and shares its close relationship with the ribozyme core. CryoEM studies of the pre-catalytic spliceosome feature intron start G+1 (5’ss) supported by Prp8 Linker domain, although they differ on the exact interactions. A recent model from Lührmann Lab (Zhang Z. et al., 2024a,b) has intron G+1 stacked on top of the aromatic ring of Phe F1551, forming two hydrogen bonds involving G+1 N1-H and C2-NH2 with a backbone carbonyl next to peptide bond of Phe. A previous model from Shi group has the same F1551 staking and two more backbone carbonyls involved: H1563 and G1564 (Zhan et al., 2018b), as in S.c. from earlier work by the same group (Wan et al., 2016a, S.c. F1623, H1635, G1636). Their version includes H-bonds involving the intron G+1 N1-H and the backbone carbonyls of Prp8 F1551 and H1563, G+1 C6=O with a peptide bond -NH of F1551 and G+1 C2-NH2 with a backbone carbonyl of G1564. It would be interesting to see how substitutions of these amino acids in Prp8 affect 5’ss recognition, as for now they are not among splice site suppressor mutations (Granger and Beggs, 2005; Galej et al., 2014). These 5’ss – Prp8 interactions were modelled in the pre-catalytic spliceosome at the transition from pre-B to B complex (Figure 5). The base pair between the ends of the intron appeared in CryoEM reconstructions only at the post-catalytic stage (P and ILS – Wilkinson et al., 2017; Bai et al. 2017; Fica et al., 2019; Zhang X. et al., 2019) and was later inferred in catalytic Step 2 spliceosome (Wilkinson et al., 2021 - C*, C Step 2 in Figure 5). 3’ss and the 3’exon were absent in CryoEM structures of the spliceosome until catalytic Step 2 stage. In contrast to this, a recent CryoEM model of Group IIA Ll.LtrB intron clearly shows 3’ss engaged at pre-catalytic stage (Liu et al., 2020). The 3’ exon also appears aligned with Id3 loop next to the 5’exon before branching. This structure is backed up by a scrupulous biochemical study of the pre- and post-catalytic Ll.LtrB ribozyme, indicating that all tertiary interactions are in place at the pre-catalytic stage (including λ-λ’ and γ-γ’, to which we devoted some attention here). There cannot be any doubt that the Ll.LtrB intron RNP particles were isolated in their native state from their natural host, as the results are consistent with in vivo biochemical probing of the Ll.LtrB ribozyme (Dong et al., 2018).
As we reviewed above, spliceosome assembly involves 3’ss from the start. It is selected and juxtaposed with the BP site in early complex E (see above 2.2 3’ss selection by protein co-factors SF1 and U2AF65/35), while the BP helix is assembled and 3’ss and 5’ss are matched together in complex A (see above 2.3 Branchpoint pairing with U2 snRNA, and 2.4 5’ and 3’ss coupling). Remarkably, both 5’ss and 3’ss are freed at the pre-catalytic stage (Figure 5, Pre-B to B transition): U1 snRNA unwinds from the 5’ss (Charenton et al., 2019) and U2AF heterodimer disassociates from the 3’ss, triggered by the phosphorylation of BP-bound SF3B1 (Kirchhoff et al., 2026). As 5’ss is unpaired and 3’ss is disengaged at the same time and in proximity to each other, there is an opportunity for them to pair.
Intron start G+1 is an invariant information-heavy position. Its recognition is unlikely to be delegated solely to a protein relying on H-bonds with the peptide backbone. Nucleic acids command information flow by specific base-pairing – as in translation by the ribosome, the older RNP-ribozyme. While the ribosome decodes RNA messages by matching codons with amino acids, the spliceosome compiles information into alternative RNA messages for the ribosome. Splicing combines flexibility, precision and speed: the spliceosome needs to make ‘sense’ (preserving the translation reading frame) moving fast. Mistakes (non-productive splicing) lead to mRNA deficiency slowing down translation and can overload nonsense-mediated decay. Fidelity is essential: information at splice sites is read by base-pairing. The intron ends pair is non-canonical, but strictly specific for the mutual recognition of the intron G+1 with G-1 (or A+1 with C-1). Besides, the intron ends pair needs to be formed at the precatalytic stage, as it is a tertiary link that repositions intermediates for the second step using the momentum of the BP helix toggle produced by the first step.
This disparity on 3’ss involvement in the pre-catalytic spliceosome we leave to the reader’s judgement, in the hope that the ‘ribozyme first’ and ‘protein first’ views can be reconciled before long. In our opinion, early involvement of the 3’ss is the only mechanism that can explain how the ancestor of spliceosmal introns thrived duplicating protosplice sites and how modern splicesomal introns can operate with protosplice site repeats and join the exons with utmost precision (in Figure 5 RNA outline involves 3’ss at pre-catalytic stage).

3.3. Exon Recognition by U5 snRNA Loop1

It is best to flag the differences in the current models of U5 base-pairing with exons to orientate the reader. Since 2016 CryoEM produced six different binding registers for 5’ exon on U5 Loop1; the loop was modelled as 11 and 7nt (five schematics outlined in supplementary materials of Artemyeva-Isman and Porter, 2021 and the 6th with bulged exon position -3 reported in Zhang W. et al., 2024). The current CryoEM models feature a 7nt loop, placing the 5’ exon end at U5 Loop1 U40. C39C38 remain free to bind the 3’ exon, for which base-pairing is not specified. The difficulties in placing the exons on U5 Loop1 might be due to relying too much on yeast genetics. Initially, the spliceosome was studied in S. cerevisiae. The distinct evolutionary trajectories between S.c. and humans (see Figure 4C in Part 1 - Artemyeva-Isman, 2026a) led to markedly different conservation patterns of intron and exon splice sites in these species. While U6 and U2 binding sites in S.c. introns are near-perfect, the exon ends are not conserved, so aligning with U5 Loop1 is impossible. This led to the initial impression that U5 Loop1 interactions with the exons are not as important as that of U6 and U2 with the intron. Early crosslinking experiments in S.c. (Newman et al., 1995) and HeLa nuclear extracts (Sontheimer and Steitz, 1993) involved substitutions of the guanines of the exon ends for 4-thio-uridines, which did not help to elucidate the binding register of the lost guanines and predicted non-Watson-Crick pairs for the exon ends. Strangely, the human consensus for the exon ends was not accounted for in the effort to place the exons on U5 Loop1. The current CryoEM model presumes that 82% conserved G at the end of human exons forms a wobble G•U pair with U5 Loop1 U40. Why a G•U wobble appears the most conserved pair is not clear – Group IIA introns use Watson-Crick pairs for exon recognition. Moreover, this is not a standard G•U wobble, which could have been an acceptable replacement for a Watson-Crick pair, but a bifurcated G•U (see Zhang Z. et al., 2024a,b); O4 of U makes bifurcated H-bonds with N1-H and N2-H2 of G. This requires U to turn ~450 compared to the standard G•U wobble. ‘Such base pairs [bifurcated G•U] are not well accommodated within a regular helix, and, instead, they occur close to other types of non-Watson–Crick pairs (e.g., trans Hoogsteen)’ (Masquida and Westhof, 2000). This is hard to understand, as Watson-Crick pairs are still allowed to recognise other exon positions under the CryoEM model. The emphasis is put on protein recognition: Prp8 Lys K1306 forms an H-bond with N7 of the exon-end G. Zhang Z. et al., 2024a still refer to Ségault et al., 1999 study of U5 mutants in HeLa nuclear extracts where the conserved Loop1 appeared dispensable for both steps of splicing, ‘indicating that PRP8 can functionally compensate for U5 loop 1 nts during mammalian splicing’. This extreme view cannot be reconciled with a growing body of genetic evidence, particularly concerning severe neurodevelopmental abnormalities linked to U5 Loop1 mutations in humans (see below).
In 2021 we proposed a new U5 model based on homology to Group IIA Id3 loop and statistical analyses of positional dependencies at human splice sites (Artemyeva-Isman and Porter, 2021). U5 Loop1 in open conformation is 11nt and the exons meet at U5 C39|C38. This model provides maximum possible Watson-Crick pairs for human exons and U5 Loop1 with G=C pairs for the conserved guanines of the exon ends. Contrary to the CryoEM reconstruction, our model can explain the new genetic data by specific U5 base-pairing with the exons (see below).
Homology to Group IIA Id3 loop is our starting point. According to the recent CryoEM structure (Liu et al., 2020) at the pre-catalytic stage folding of the Group IIA Ll.LtrB intron is already complete, with the 5’ss helix, the catalytic site, the BP helix and all the tertiary links in place. Notably, 3’ss is close to 5’ss, and 3’ exon is aligned with Id3 Loop alongside the 5’exon before the branching reaction. Id3 Loop pairs the last 7nt of the 5’ exon and the first 4nt of the 3’ exon (Figure 2B in Part 1 - Artemyeva-Isman, 2026a) (Ichiyanagi et al., 2002). In spliceosomal introns U5 Loop1 also binds the exons asymmetrically, although there is a shift of the binding register by 1nt: last 8nt of the 5’ exon and the first 3nt of the 3’ exon (Figure 2A; Artemyeva-Isman and Porter, 2021). This coincides with the shift of the intron ends pair by 1nt at the 3’ss.
We proposed a mechanism that explains how universal U5 Loop1 pairs the multitude of diverse exons (see Part 1 - Artemyeva-Isman, 2026a; Artemyeva-Isman and Porter, 2021). This mechanism is rooted in Group IIA intron mobility. When choosing new targets in the genome, the Ll.LtrB intron seeks the best approximation to the shape of the helix of all Watson-Crick pairs between Id3 Loop and the target site, as with its exons. However, almost half of these base-pairs are not true Watson-Crick, but C--U, U--U, G--U, and C--A – all of which were previously reported in crystal structures of different biological systems as isosteric to canonical, or mimic pairs assuming a Watson-Crick-like shape (Bebenek et al., 2011; Wang et al., 2011; Rozov et al., 2015; Rypniewski et al., 2016). This can be achieved by tautomerisation - migration of a proton – in one of the bases of these pairs (Artemyeva-Isman and Porter, 2021). On the contrary, A••A, A••G, G••G and C••C that cannot mimic Watson-Crick geometry are kept out of these interactions. We suggested that the optimal helix architecture is attained by a combination of Watson-Crick and mimic isosteric pairs. The same types of base-pairs are typical in U5 Loop1 duplexes with human exon ends. Base tautomerisation is not unusual in active sites of ribozymes and riboswitches. Now I suggest that tautomeric bases making isosteric pairs routinely support exon recognition and allow for diversification of exon sequences to promote protein evolution (see Part 1 - Artemyeva-Isman, 2026a). This said, isosteric pairs are merely tolerated: Watson-Crick pairs still command the recognition. They are most conserved proximal to exon ends (Figure 2C) and their mutations lead to aberrant splicing and disease (Fu et al., 2011; Ohno et al., 2005; Baeza-Centurion et al., 2020 and references therein; Lee et al., 2025; Artemyeva-Isman and Porter, 2021).
In support of 3’ss involvement at pre-catalytic stage I created a U5 model that includes the intron ends’ pair (upgraded SimRNAweb v2.0 tool Moafinejad et al., 2024, Figure 7). In 3D U5 Loop1 appears sandwiched by the exons, with the intron termini pair wedged between the 5’ and 3’ss (Figure 7A). Exons are paired to U5 Loop1 side by side; both bonds with the intron are intact (5’ and 3’ss indicated by black arrows). U5 Loop1 strand bends sharply between C38 and C39, where base-pairing switches from one exon to the other (Figure 7B,D). It follows a smooth curve with the longer 5’exon duplex. Convincingly, U5 strand is less flexible here because of two pseudouridines: their water bridges confer rigidity to the sugar-phosphate backbone from Um41 to Ψ43 and from A44 to Ψ46 (Figure 7D, properties of pseudouridine explained above in 2.2 3’ss selection by protein co-factors SF1 and U2AF65/35). For this model I used hypothetical exons that form all Watson-Crick pairs with U5. These exons represent an enhanced consensus, which applies when intron splice sites are weak (Figure 7D shows 5’ exon consensus for introns that lack conserved intron G+5). The idea is that all the exons approximate to this model using isosteric configurations for mismatched pairs.
Together the exons and U5 Loop1 make up the 11bp turn of the RNA A-helix, characteristically hollow in the middle (Figure 7C). The future junction of exons, i.e. mRNA can be disassociated from U5 Loop1 without needing to unwind the stem: the loop fits inside a helical turn while the exons form the outside strand (Figure 7A, admittedly this structure is an improvement on our previous attempt, where mRNA went through the loop – Artemyeva-Isman and Porter, 2021). However, it is obvious that the size of a recognition loop cannot exceed 11nt to avoid entangled strands and the need for topoisomerase activity to release joined exons after splicing. Extending recognition by 6 more nucleotides in the 5’ exon requires a separate Id1 loop in Group IIA introns (Figure 7B).
The intron termini guanines pair with Watson-Crick edges, their glycosidic bonds in trans orientation (Figure 7A,B and see previous section). This tWW pair implies parallel strands orientation, which is not obvious in the presented structure, except for the RNA strand flip visible at the 3’ss. At first glance, we found this disappointing: instead of a slight flip at 3’ss, no dramatic feature appeared at 5’ss to produce a change in the direction of RNA strands. CryoEM structure of the pre-catalytic Ll.LtrB Group IIA intron shows that ‘the 5′ end of the intron (near G1) forms a notable kink, which possibly facilitates the initiation of branching’ (Liu et al., 2020). However, this kink appears 3’ of the intron G+1, which is outside this predicted structure - the shape of RNA backbone at the 5’ss will require modelling the intron strand. Their CryoEM Group IIA structure also shows a clear bend exactly at the 3’ss. Perhaps the flip of the RNA strand at the 3’ss, chosen by SimRNA algorithm for this structure reflects the real thing: a very short duplex is more easily amenable to such a change. This explains the conserved asymmetry of exons binding the recognition loop: Group IIA Id3 loop forms 7 and 4 bp with 5’ and 3’ exons, while spliceosomal U5 Loop1 exaggerates the ratio to 8 and 3 bp. Notably, spliceosomal introns can accommodate G+1••G-1 pair, while Group IIA introns have it staggered by one base G+1••A-2 – is it to release tension, because the 3’exon duplex is 4bp instead of 3bp? In any case, splice sites are distinct from the helical continuity of the exons with U5 Loop 1, where RNA pairs neatly stack on each other. A gap demarcates the intron ends pair from the exons as the intron comes off U5 Loop1 (Figure 7C projection).
Returning to the three questions on lariat ribozymes (Q1-3, previous section), we suggest the following explanations: (Q1) Spliceosomal introns have a different intron ends pair, due to protosplice site repeats at both 5’ss and 3’ss. This pair in tWW configuration safeguards U5 Loop1 from binding intron ends. Intron G+1 pairing with G-1 instead of A-2 likely provoked the shortening of the 3’exon interaction with U5 Loop1 by 1bp, compared to Group IIA introns. (Q2) Exons bind the recognition loop asymmetrically to accommodate intron ends pair with parallel local strands orientation. The short 3’ exon duplex allows the RNA strand to flip. (Q3) Splice sites are structurally defined by a gap between the distinct intron ends pair and U5 Loop1 helices with the exons, where RNA pairs are stacked (Figure 7C). The ribozyme cleaves the sites where the intron strand comes off the loop.
Now let us examine how our U5 model performs. Firstly, it clarifies dependencies between conserved positions at human splice sites (Artemyeva-Isman and Porter, 2021). A stable interaction of U6 with the start of the intron – 5’ss helix - depends greatly on a C=G pair between U6 C42 and G+5 conserved in 78% of human introns. The remaining introns which lack G+5 are preceded by 5’ exons that can form significantly more Watson-Crick pairs with U5 snRNA at exon positions -1, -2, -3 and -5 (numbering always from the splice site). Specifically, we link this effect to U5 rather than U1 snRNA, because of the choice for adenine, not cytosine in exon position -3 and the effect of exon position -5, where U1 base-pairing does not extend. In turn, human exons conserve the last G at 82%. Reciprocally, exons lacking the last G are followed by introns that form significantly more Watson-Crick pairs with U6 at intron positions +5, +6, +7 and +8. The conclusion is that U5 and U6 snRNAs recognise the exon-intron boundary collectively, mutually compensating for the variability of their binding sites. Accordingly, a variety of recent studies (reviewed below) confirmed the compensatory role of U5 base-pairing with the exons when U6 binding to the intron is weak or undermined.
A phylogenomic study (Olthof et al., 2024) of 263 eukaryotic genomes investigated the evolution of minor introns. The general trend is confirmed to be minor to major intron conversion with minor-like introns identified as an intermediate state. As a reminder, minor and major classes are distinguished by U11 and U6atac recognising intron C+5C+6, as opposed to U1 and U6 pairing with G+5 and minor-like set has U6atac-type, not U6-type 5’ss. However, RNA-seq datasets for deficiencies in minor spliceosome components show that ~86% of human minor-like introns are still spliced out. A closer investigation revealed a marked difference in exon logos between the remainder of minor-like introns that were retained and the introns that were removed: exon end G-1 appeared to be much more frequent in spliced out introns. The same, but to a smaller extent was true about exon start G+1. In addition, they found that these spliced-out minor-like introns were bound by U2AF – the major spliceosome component - under minor spliceosome suppression. The authors conclude that major spliceosome competes with the minor spliceosome and can process minor-like introns preceded by exons that end with the conserved G. This points at specific U5 base-pairing as compensation for the low affinity of U6 to U6atac binding site at the start of the intron. To be precise, ‘in the absence of the [exon] -1G there are insufficient Watson-Crick interactions to support splicing of the minor-like introns by the major spliceosome’.
Recent studies of mutations that impair U6 snRNA binding or 5’ss helix stability observed the enhanced role of U5 in exon recognition. METTL16methylates N6 of U6 Am643 (human nt numbering, Figure 2A in Part 1 - Artemyeva-Isman, 2026a); in vertebrates it also controls methylations in general by targeting an ACAGAGA hairpin in the 3’UTR of MAT2A pre-mRNA. (MAT2A encodes S-adenosylmethionine (SAM) synthetase. SAM is a methyl donor for nearly all cellular methylations2 Pendleton et al., 2017.) Mettl16 KO in yeast S. pombe (Ishigami et al., 2021) and a splice site mutation in the homologous FIO1 gene linked to cold-flowering Arabidopsis thaliana (Parker et al., 2022) display globally altered splicing patterns. 5’ss usage declines in the absence of a strong 5’ exon consensus A-3A-2G-1 and intron U+4 is favoured over the usual A+4. Intron U+4 is explained by a Watson-Crick U-A pair, since a methyl-N6 is not there to obstruct it (Figure 2A,C). Enhanced exon consensus is explained by U5 base-pairing that supports U6 interaction with the intron – the 5’ss helix destabilised by the absence of Am6 stacking (see 3.1 5’ss helix). In their following study Parker et al., 2025 separate two classes of human 5’ss: ‘defined by preferential base-pairing potentials with either U5 snRNA loop 1 or the U6 snRNA ACAGA box’ but relating to the CryoEM U5 base-pairing model with the exon-end G-1•U40 pair. Lamentably, the superior energy benefit of the G-1=C39 pair is lost for the 5’exon here. Moreover, the authors mention our computational analyses (Artemyeva-Isman and Porter, 2021) as showing that compensation for weak intronic positions at 5’ss is better explained by U5, rather than U1 Watson-Crick base-pairing. In this case we meant U5 base-pairing according to our model. To precise, 5’ exon A-5 is significantly on the rise, when intron lacks G+5. This makes A-5–Ψ43 pair in our U5 model, but in the CryoEM model A-5 aligns with U5 A44, which does not supply a Watson-Crick pair. Nevertheless, it is gratifying that ‘the cooperative and compensatory activity of U5 and U6 snRNA interactions at 5′ splice sites’ is now an accepted fact.
A very different gene, but the same U5 effect is linked to newly identified ReNU neurodevelopmental syndrome (Greene et al., 2024; Chen et al., 2024; Nava et al., 2025). RNU4-2 is one of the two human U4 snRNA genes. RNU4-2 has a more important expression in the brain than the RNU4-1 copy. U4 snRNA extensively binds U6 to prevent premature catalytic site folding. U6 nt30-36 are engaged in U6/U4 Stem III; U6 nt37-46 with ACAGAG are looped out and secured downstream by the U4 quasi-pseudoknot and Stem I, engaging U6 nt47-56 (Charenton et al., 2019, Figure 6A, explained above under 3 Pre-catalytic spliceosome: fidelity check). Stem III unwinding is triggered by U6 ACAGAGA pairing the intron. The range of U4 involved in stem III will be the loading site for Brr2 helicase. In turn, when Brr2 unwinds stems I and II, U6 stretch engaged in stem I configures U2/U6 helix I and the catalytic triple helix, and U6 range involved in stem II folds into U6 ISL (Figure 2A). Pathogenic U4 mutations disrupt Stem III, the pseudoknot and the unpaired residues between them, where RBM42 binds to support the RNA conformation. Splicing abnormalities picked up in patients’ blood samples show suppressed intron and enhanced exon consensus at 5’ss (Nava et al., 2025). The conclusions again pointed at cooperative, or rather, collective U5 and U6 5’ss recognition. U4 mutations interfere with normal U6 binding to the intron start, possibly by distorting the conformation of the ACAGAG loop of the U4/U6 complex. U5 snRNA compensates for U6 binding deficiency by forming a stronger duplex with the end of the exon.
The same mechanism of compensation by U5 is likely at work at the 3’ss. Returning to the positional dependencies at human splice sites (Artemyeva-Isman and Porter, 2021) intron position -3 is conserved as 65% C and 30% U, and suggested to pair with U2 G31 (see above 2.3 Branchpoint pairing with U2 snRNA, and 2.4 5’ and 3’ss coupling). Introns that lack C-3 are followed by 3’ exons that start significantly more often with G+1. And finally, an observation concerning the 3’exon itself. Human mutations of the exon-start G+1 have very different effects on exon inclusion: from complete suppression to no influence at all. Exon C+2 and G+3 appear to compensate for G+1 substitutions – by providing C=G/ G=C pairs with U5 Loop1 Gm37 and C36. The general conclusions are that (1) the optimal U5 binding register is achieved by exons meeting at U5 C39|C38 – the maximum possible Watson-Crick pairs including the most conserved exon-end guanine. (2) U5 snRNA specific base-pairing with the exons compensates for weak intronic splice sites both at 5’ and 3’ss to ensure splicing precision. (3) Any of the three G=C/ C=G pairs can stabilise the 3’exon/U5 duplex (Figure 2A, Figure 7).
Mutation studies published in 2025 allude even more pointedly to U5 Loop1 base-pairing with the exons. Prp16/DHX38 ATPase is bound to the spliceosome from the pre-catalytic stage - complex B. Prp16 acts between the two catalytic steps clearing the first step proteins Cwc25, Yju2 and Isy1 to be replaced by the second step proteins (see Figure 5). Human DHX38 is one of a ~100 genes linked to retinitis pigmentosa (Obuća et al., 2022), the most common inherited retinal degeneration, which is caused by mutations in retina-specific proteins and ubiquitous splicing factors affecting retina-specific transcripts, including Prp8, Brr2 and other protein components of the U5•U6/U4 tri-snRNP (Yang et al., 2022). Prp16 appears to be responsible for diversifying exon recognition in pathogenic yeast Cryptococcus neoformans. Conditional Prp16-KO mutants display a marked growth delay and suffer a global increase in intron retention under non-permissive conditions. A small fraction of introns (7%, 2808 of ~40000 total) escapes this fate (Negi et al., 2025). These unaffected introns significantly differ from the rest in their exon logos: their 5’exon ends with A-3A-2G-1 and the 3’exon starts with G+1, with both guanines especially on the rise. Moreover, dominant Prp16 mutations are known to suppress splicing in C. cerevisiae and do likewise in C. neoformans – except for these unaffected introns flanked by well conserved exon ends. The authors conclude that C.n. Prp16 role is to support suboptimal exon interactions with U5 and that exons that can form Watson-Crick pairs with U5 Loop1, especially conserving guanines at their ends, render Prp16 support unnecessary.
Finally, mutations in U5 snRNA Loop1 itself are discovered to cause severe neurodevelopmental diseases NDD (Jackson et al., 2025; Nava et al., 2025). Human U5 snRNA is encoded by five distinct variant genes, all U5 variants have identical Loop1 sequence. U5 variants predominantly included in the spliceosome are RNU5B-1 and 5A-1, the Loop1 motif of these two genes is deplete of variation by strong purifying selection. To be precise, gnomAD v4.1 (n>800,000) records 14 C T changes in RNU5B-1 Loop1 motif: 2 C36, 1 C38, 3 C39 and 8 C45 and 11 C T changes in RNU5A-1 Loop1 motif: 4 C36, 2 C38, 4 C39 and 1 C45. Other single nt variations SNVs involve only the two distal positions of Loop1 C45 and T46: 2 in RNU5B-1 and 4 in RNU5A-1. In comparison gnomAD v4.1 records 550, 642 and 5876 various changes in the U5 Loop1 motif of RNU5E-1, F-1 and D-1 genes, respectively. Evidently, C U and any other SNVs of C45 and U46 in Loop1 are very rare in spliceosomal U5 and are found in heterozygous state for only one of the two predominant U5 variants. Their clinical significance is unknown. Analyses of seven databases of rare diseases (Jackson et al., 2025; Nava et al., 2025) yielded seven pathogenic U5 Loop1 dominant mutations in 19 individuals with NDD (mainly in 100KGP and PFMG20253) – 16 mutations in RNU5B-1 and 3 in RNU5A-1, all distinct from SNVs recorded in gnomAD. These cases are rare suggesting that pathogenic U5 Loop1 mutations of the predominant U5 variants are more frequently lethal before birth. These seven U5 mutations seem to have individual splicing signatures in patients’ blood samples (PCA principal components analysis, Nava et al., 2025). The most frequent mutation, 7/19, 1 in RNU5A (Nava et al., 2025; Jackson et al., 2025 – data overlaps for UK cases) is U5 Loop1 C39 G39 substitution, which affects recognition of the 5’ss (Nava et al., 2025). In the CryoEM U5 model C39 is not interacting with the 5’exon, which is replicated by the explanatory figure presented by Nava et al., 2025. According to this model the influence of U5 C39 on 5’ss recognition cannot be clearly accounted for. On the contrary, the importance of U5 C39 is clear in our U5 model (Figure 2A and Figure 7): C39 pairs with the 5’exon end G-1 conserved in 82% of human exons. Our U5 model was used in the explanatory figure by Jackson et al., 2025. Next, two C38 U38 changes, like in gnomAD were also found in datasets examined by Nava et al., 2025, but these were changes of unknown significance, rather than proven pathogenic. The expected effect is a change from the exon-start G+1=C38 pair, conserved in 50% of human exons to an isosteric G+1--U38 /wobble G+1•U38. We have already observed that exon G+1 mutations can be compensated by the other two G=C/C=G pairs of the 3’exon (Artemyeva-Isman and Porter, 2021). Curiously, Gm37 C37, 1/19 and Gm37 U37, 1/19 are both pathogenic. Other mutations affect the size and conformation of U5 Loop1. Insertions of adenines occur in the uridine stretch of Loop1 between U40_Um41, 2/19 (in RNU5A), U4243, 3/19, and in a single case uridine is added after C39. Enlarging the loop by 1nt can perturb the binding register of the exons: U40_Um41 insA apparently affects both 5’ and 3’ss recognition (Nava et al., 2025). On the contrary U5 A44 G44 change, 4/19, leads to reduction of the size of the recognition loop. Unpaired wt U5 Loop1 is likely to be reduced to 7nt by the Gm37=Cm45 pair, but it is an unstable conformation as the next pair in the prolonged stem is the mismatched C36--Ψ46. When paired with exons, U5 Loop1 opens to its full extent of 11nt, allowing for specific 3’exon binding. U5 A44 G44 mutation stabilises the prolonged stem by a new G44=C38 pair, reducing the loop to only 5nt. U5 binding site for the 3’ exon is then unavailable - accordingly A44 G44 mutation disrupts 3’ss recognition (Nava et al., 2025).
The importance of getting the U5 base-pairing register right is not a matter of theoretical merits - on the contrary, it is essential for practical applications in medical genetics and gene therapy (see Part 4, Artemyeva-Isman, 2026d). The recent studies reviewed above led to renewed interest in U5 and the revival of experiments with Loop1 mutations, that were not attempted since the 1990s (Cortes et al., 1993). However, recent studies in yeast S. pombe (Ishigami et al., 2021) and C. neoformans (Negi et al., 2025) do not conclusively prove either one or the other binding register. Splicing activation by transiently expressed mutant U5 in S. pombe for nhm1 intron 2 and ysh1 intron 1 is better explained by the CryoEM binding register with 5’exon end at U5 U40, but trk2 intron 3 and sec71 intron 1 splicing activation in the same experiment better fits our base-paring model with the exon junction at U5 C39|C38, while atg1803 intron 2 splicing suppression can be explained by both binding registers (Ishigami et al., 2021). Mutated U5 integrated into a safe haven locus in C. neoformans showed only a non-significant splicing increase for pac1 intron 5, better explained by our U5 model with exons meeting at C39|C38, but activation of mps1 intron 6 splicing does not fit either binding register. On the contrary, the integrated pac1 minigene construct with mutations in exon 5 T-3T-2C-1   A-3A-2G-1 and exon 6 A+1 G+3 produced a robust intron 5 splicing increase in Prp16-KO cells under non-permissive conditions, consistent with exons that conserve strong U5 Loop1 binding sites and are spliced independently from Prp16 (Negi et al., 2025). These experiments show that splicing is readily activated by changing exons to match wt U5 Loop1, but improving splicing of a target intron by an additional U5 gene with a changed Loop1 motif is less predictable. Moreover, the importance of U5 interactions with exons in yeast is less significant than in humans, as yeast do not conserve exon ends - although S. pombe and C. neoformans have a weak indication of exon consensus, as opposed to S. cerevisiae, that has none. Exon interactions with the recognition loop descend from the ancestral Group IIA intron, so yeast introns are further from the ancestral ribozyme state than human introns. In S.c. this is concurrent with massive intron loss, while C.n. and S.p preserved a lot more introns, but their introns are drastically truncated (C.n. intron average is 56nt) and the splicing mechanism was adjusted for short introns (Negi et al., 2025). The eukaryotic intron progenitor was undoubtedly much longer, as most modern Group IIA introns are longer than the average length of introns in humans. These indications suggest that yeast are not good hosts for U5 experiments - possible experiments in human cells, addressing the problem of the U5 binding register, are discussed below (see Part 4, Artemyeva-Isman, 2026d). Finally, it remains to be considered if the possibility of a U5 binding register shift can be absolutely excluded for the exons that lack the end G-1.
Coming back to spliceosome assembly, exon recognition by U5 Loop1 is the finishing snRNA interaction with the pre-mRNA and therefore responsible for the final 5’ss and 3’ss definition in the pre-catalytic complex. 5’ss is defined collectively by U6 snRNA, the 5’-3’ss base pair, and U5 snRNA. The involvement of 3’ss is required for 5’ss definition, because the base pair of the intron ends reconstructs the junction of exons within the splice site repeats to ensure correct U5 binding. Moreover, this 5’-3’ss link is essential to establish before branching to connect the two catalytic steps. This link is responsible for transmitting conformational change of the 5’ss-BP helix to the 3’ss to re-position 3’exon for ligation. The intron termini pair preceding U5 binding implies that both exons are paired with Loop1 at the pre-catalytic stage like in the recent Group IIA intron CryoEM structure (Liu et al., 2020). 3’ss definition is finished collectively by U2 snRNA (provided U2 G31 pairs the conserved intron C-3, see 2.3 Branchpoint pairing with U2 snRNA, and 2.4 5’ and 3’ss coupling), the 5’-3’ss base pair, and U5 snRNA. The cleavage sites are structurally demarcated by a gap between the distinct intron ends pair and the helices of U5 Loop1 with the exons where pairs are stacked together (see Figure 7C projection). Protein factors - most obviously Prp8 and apparently Prp16 (in C. neoformans experiment Negi et al., 2025), stabilise suboptimal U5 Loop1 interactions with the exons.

3.4. Catalytic Centre Formation After Brr2 Displaces U4

The RNA core at pre-catalytic spliceosome complex B (Figure 5,6A) includes the 5’ss helix, intron ends pair, exons paired with U5 Loop1 and the BP helix, albeit bound by SF3 protein complex that prevents active conformation of the BP A. U4 snRNA is bound to U6 snRNA blocking catalytic centre formation until the definition of the 5’ss is completed by U5 snRNA. The stable 5’ss helix triggers Brr2 helicase translocation from its binding site on U4 snRNA in the 3’ to 5’ direction: unwinding U4/U6 stem I and then stem II (Nielsen and Stailey, 2012). The freed stretch of U6 undergoes complete structural rearrangement forming an internal stem-loop U6 ISL and base-pairing with U2 - U2/U6 helix I (Figure 2A). These elements then assemble a triple helix of the tight catalytic centre that coordinates two Mg2+ ions. NTC/NTR protein complex (‘nineteen complex’ - Prp19 and associated proteins) is recruited to support the configuration of the U6/U2 catalytic core - CryoEM studies (Wilkinson et al., 2020) define this stage as Bact complex (Figure 5).

4. Catalytic Spliceosome: Reversible Intron Splicing

4.1. BP Helix Preparation for Branching

The last step in spliceosome activation is the displacement of the SF3 protein complex by Prp2/DHX16 (Figure 5). This allows for the precise coordination of the BP A provided by hydrogen bonds with the conserved intron residues UBP-2 and A-2 (Figure 2A in Part 1, Artemyeva-Isman, 2026a) captured by CryoEM studies (Wilkinson et al., 2021). The Watson-Crick edge of BP A pairs with the sugar edge of UBP-2, glycosidic bonds are in cis orientation, cWS pair (see Leontis et al., 2002 for geometric classification of RNA base pairs). Homologous interaction is reported in Group IIB introns with the conserved CBP-2 (Xu et al., 2023) and presumed in Group IIA Ll.LtrB (Figure 2B). When the Hoogsteen edge of BP A pairs with the Hoogsteen edge of the invariant A-2, glycosidic bonds are in trans orientation, tHH pair (Wilkinson et al., 2020). Although homologous interaction was not yet shown in Group II introns, the same pairing is possible as A-2 Watson-Crick edge is involved in the intron ends pair (Figure 2B), whereas the Hoogsteen edge is free to interact with the BP A.

4.2. Two Mg2+ Ribozyme Catalysis and Its Reversibility

Splicing catalysis is a 2-step process (see Part 1, Artemyeva-Isman, 2026a), involving two reactions that break and rejoin the RNA sugar-phosphate backbone. Both splicing reactions are catalysed by the same pair of catalytic Mg2+ ions, coordinated by the U2/U6 catalytic centre (Figure 3A,B in Part 1, Artemyeva-Isman, 2026a). The first splicing reaction is initiated by the 2’O of the ribose of the BP A – the sugar edge of the BP A is turned outwards – while the Watson-Crick and Hoogsteen edges are clamped by UBP-2 and A-2 as described above. The first step protein factors Cwc25, Yju2 and Isy1 further support the correct ribozyme conformation. The 2’O of the ribose of the BP A is activated by one of the Mg2+ (M2, Figure 3A) and launches a nucleophilic attack on the 3’O of the ribose of the last nucleotide of the 5’exon – usually a G, that is stabilised by the other Mg2+ ion (M1, Figure 3A). Both Mg2+ ions coordinate the intermediate pentavalent state of the phosphate: the scissile bond with the 3’O of the exon-end G and the emerging bond with the 2’O of the BP A. This is the RNA chemistry of the Catalytic Step 1 spliceosome (designated as complex B* in CryoEM studies, Wilkinson et al., 2020, Figure 5).
Next, is the transition stage: the ribozyme re-arranges the intermediates and prepares for the second reaction. The intron is now shaped like a lariat – the intron 5’ end is linked via its phosphate to the 2’O of BP A. Indeed, the ‘branched’ BP A ribose has 2’, 3’ and 5’ bonds with three different phosphates (Figure 3B). The conformational toggle upon the formation of this bulky lariat intermediate drives the BP A and the intron 5’ end out of the catalytic centre. The intron ends are joined together by a base pair - this link tows on to the 3’ ss – the intron 3’ end and the 3’ exon slide into the positions vacated by BP A and the intron 5’ end. The precision of the stereochemistry of this transition may be the reason for unchanged lariat ribozyme structure persisting in Group II introns and spliceosomal introns (Figure 2A,B) through billions of years (Figure 4 in Part 1, Artemyeva-Isman, 2026a).
Prp16 displaces the first step protein factors and the second step factors join the complex - Figure 5. The catalytic Mg2+ ions reverse their roles (Figure 3A,B): now the 3’O of the ribose of the 5’exon is activated by the 1st Mg2+ (M1, Figure 3B) and becomes a nucleophile attacking the 3’O of the ribose of the last guanine G-1 of the intron, which is in turn supported by the 2nd Mg2+ (M2, Figure 3B). The fact that catalytic metal ions can act both for activation of bonding groups and support of leaving groups is at the basis of reversibility of splicing catalysis. Again, both Mg2+ ions coordinate the intermediate pentavalent state of the phosphate: the scissile bond with the 3’O of the intron 3’end G-1 and the emerging bond with the 3’O of the 5’exon-end G. 5’ exon stays in the same position throughout both steps of splicing. The catalytic Step 2 spliceosome complex (Wilkinson et al., 2020, Figure 5) was designated as complex C in CryoEM studies.
Both exons are bound to U5 Loop1 when ligation occurs - the junction of the exons initially stays bound to U5 (this stage was designated as post-catalytic complex P) and gets disassociated with the help of Prp22/DHX8 (Figure 5). U5 Loop1 without the RNA partner is likely to collapse to 7nt ‘closed’ conformation with Gm37=Cm45 prolonging the stem, despite the C36--Ψ46 mismatch. The remaining complex of the excised intron lariat, U6/U2 and U5 snRNPs stays together and can be readily isolated. Moreover, this complex - known as intron lariat spliceosome ILS is stable until the lariat is debranched by DBR1, a unique enzyme that can cleave 2’O ribose bond. Dbr1-/- is embryonic lethal in mice, and a recent DBR1-/- 293T cell line accumulates ILS, leading to deficit of U5, U6 and U2 snRNPs and widespread increase in exon skipping (Buerer et al., 2024). Notably, DBR1 activity is sequence-specific: lariats of conserved BP and 5’ss sites are debranched faster, than their more divergent counterparts.
Most relevant for practical applications is the homology of ILS to the Group IIA intron RNP particle – the mobile form of the intron in search of a suitable genomic target with a fully assembled catalytic core for reverse splicing. Eukaryotic pre-mRNA splicing also works in reverse – as was shown experimentally in S. cerevisiae (see Part 1, Artemyeva-Isman, 2026a; Tseng and Cheng 2008; Lee and Stevens 2016). Reverse splicing is suggested to be implicated in splicing quality control (Smith and Konarska, 2008). Genomic integration of introns by ILS and its possible application for gene repair will be discussed in Part 4, Artemyeva-Isman, 2026d.

5. Discussion

Every pre-mRNA splicing event is preceded by a highly coordinated assembly of the ribozyme with RNA elements arriving in complex with their protein co-factors as snRNPs.
Human spliceosome includes more than 150 proteins: the RNA component of snRNPs is relatively small. The fragmented ribozyme assembled in trans has lost ancestral scaffold and allosteric co-factor functions (Group IIA intron domains I and II reviewed in Part 1, Artemyeva-Isman, 2026a) and Prp8 and other spliceosomal proteins assumed these roles. snRNAs are protected, processed and chaperoned by their protein co-factors; however, RNA retains primary control. The ribozyme recognises cleavage sites by base-pairing and 3D folding. Typically, proteins are not trusted even with the new eukaryotic functions of alternative splice site choices and the need to control splicing fidelity: the two new recruits, U1 and U4 snRNAs (see Part 1, Artemyeva-Isman, 2026a) still use base-pairing to support splicing flexibility and ensure precision. Focusing on the RNA motifs and ribozyme folding we challenge the assembly narrative of spliceosomal CryoEM studies. It is established that in the early spliceosome U2AF65 juxtaposes BP and the 3’ss and U1 and U2 snRNPs cooperate to select a 5’ss – 3’ss pair. However, at pre-catalytic stage spliceosomal CryoEM structures lack 3’ss. We argue that the ribozyme structure must include 3’ss as featured by CryoEM for pre-catalytic Group IIA intron. Conservation of splice sites preserves the ribozyme structure and cannot be ignored by structural studies. The first question we must ask is how spliceosome manages to splice within CAG|GU protosplice site repeats that persist at both human 5’ss and 3’ss today. Reconstructing the splice junction correctly requires the involvement of the 3’ss and the ribozyme has a structural provision for this: the conserved intron ends pair, forming the tertiary link between G+1 and G-1. The exact configuration of this pair is tWW as established by mutation analyses, not cWH as modelled by CryoEM, as the participation of G-1 exocyclic C2-NH2 group was excluded by nucleotide analogue experiments and because cWH is impossible for the only covariant intron termini combination A+1...C-1 in 1% of human introns. The intron termini pair is also a mechanistic device used by the lariat ribozyme to connect 5’ss and 3’ss for the transition between the two steps of splicing. The first reaction induces conformational change: the lariat intermediate rotates, shifting the BP A and the 5’ss out of the catalytic core. Because of the tertiary link between the intron ends, 3’ss and the 3’exon slide in place of the first step reactants. The marvellous precision of this mechanism combined with ribozyme self-control in target recognition might explain the preeminent success of lariat introns in cellular life evolution. We offer a predicted structure of the exons paired with U5 Loop1 and the intron termini pair at the pre-catalytic stage. The binding register for the exons on U5 Loop1 is another contested point with CryoEM. Between 2016 and 2024 CryoEM studies offered 6 different ways of 5’exon pairing with U5, but no exact base-pairing was modelled for the 3’exon. Sadly, none of these explain why 82% of human exons end with a G. CryoEM models converge in the view that the exon-end G forms a G•U wobble with U5 Loop1 U40. Why a G•U wobble appears the most conserved pair is not clear – Group IIA introns use Watson-Crick pairs for exon recognition. Plus, the exact configuration was modelled as bifurcated G•U: this does not typically occur next to Watson-Crick pairs, which were still modelled for exon adenines -2 and -3 and U5 Loop1 uridines 41 and 42. Instead our model places the junction of exons at U5 C38|C39 with the maximum possible Watson-Crick pairs, including G=C pairs for the conserved exon ends. This structure is based on positional dependencies at human splice sites in our earlier paper, recently used to explain various phylogenetic and mutation analyses. Our model also fits the effects of human U5 Loop1 mutations on splicing reported in 2025 for neurodevelopmental diseases. The structure shows clear conformational demarcation of splice sites by a gap between the helical continuity of the exons with U5 Loop1 and the distinct intron termini pair, explaining how the cleavage sites are defined by the ribozyme. Further verification of the U5 binding register with the exons is considered in Part 4 (Artemyeva-Isman, 2026d).
A gap in knowledge worth mentioning is the absence of explanation for the conservation of intron C-3 at the 3’ss. We examine a combination of structural, biochemical, phylogenetic and mutation evidence and conclude that it is best explained by a base-pair with U2 G31 (U12 G16, see Figure 2A in Part 1 - Artemyeva-Isman, 2026a). BP helix bridging across a variable spacer with a polypyrimidine tract fixes the BP A and 3’ss at a 6nt distance. This provides a solution for the compact ribozyme structure and loops out the protein binding domain of the spacer. Part 4 discusses experimental verification of the proposed U6 G31 pair with the conserved intron C-3.
Another as yet unresolved question is the function of the conserved adenine at the start of introns in position +3 (and +4 in humans and other complex organisms). By analogy with Group II intron λ-λ’ tertiary interaction, we suggest that the conserved intron A+3 inserts into the minor groove of U6 ISL Cm63=G71 pair, anchoring 5’ss to the catalytic site. Mutation studies and the pattern of conservation suggest that A+4 can be functional in place of A+3.
The real purpose of our foray into the evolution and mechanism of pre-mRNA splicing is to supply broader knowledge and new ideas for translational research and industrial development of better therapies. Spliceosomal snRNAs are short and versatile for therapeutic delivery. Parts 3 and 4 (Artemyeva-Isman, 2026c,d) consider how snRNAs targeting stepwise modular assembly of human intron ribozymes can control or correct gene expression and why snRNAs will be better than the existing splicing and gene therapies.

Acknowledgments

I am grateful to Dr S. Naeim Moafinejad and Professor Janusz Bujnicki (IIMCB, Warsaw) for their contribution to building the improved U5 model: their upgraded SimRNAweb v2.0 algorithm allows for an input of base pairs defined by the Westhof geometric classification for multiple RNA molecules (Moafinejad et al., 2024). I thank Dr Clément Charenton (IGBMC, Strasbourg) for a communication that clarified some features of the pre-catalytic spliceosome structure.

Data Availability Statement

The input file for the SimRNAweb and the result pdb file used to generate the projections of the U5 structure in Figure 7 can be found in the GitHub repository https://github.com/oartem02/U5-snRNA-role-in-splice-site-definition-3D-model.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Almada, A. E., Wu, X., Kriz, A. J., Burge, C. B., & Sharp, P. A. (2013). Promoter directionality is controlled by U1 snRNP and polyadenylation signals. Nature, 499(7458), 360–363. [CrossRef]
  2. Artemyeva-Isman, O. V. (2026a). The origins of spliceosomal ribozymes. Preprints.org.
  3. Artemyeva-Isman, O. V. (2026c). U1 (and U7) snRNA for splicing therapy. Preprints.org.
  4. Artemyeva-Isman, O. V. (2026d). U2, U6 and U5 snRNAs for splicing and gene therapy. Preprints.org.
  5. Artemyeva-Isman, O. V., & Porter, A. C. G. (2021). U5 snRNA Interactions With Exons Ensure Splicing Precision. Frontiers in genetics, 12, 676971. [CrossRef]
  6. Baeza-Centurion, P., Miñana, B., Valcárcel, J., & Lehner, B. (2020). Mutations primarily alter the inclusion of alternatively spliced exons. eLife, 9, e59959. [CrossRef]
  7. Bai, R., Yan, C., Wan, R., Lei, J., & Shi, Y. (2017). Structure of the Post-catalytic Spliceosome from Saccharomyces cerevisiae. Cell, 171(7), 1589–1598.e8. [CrossRef]
  8. Battle, D. J., & Doudna, J. A. (2002). Specificity of RNA-RNA helix recognition. Proceedings of the National Academy of Sciences of the United States of America, 99(18), 11676–11681. [CrossRef]
  9. Bebenek, K., Pedersen, L. C., & Kunkel, T. A. (2011). Replication infidelity via a mismatch with Watson-Crick geometry. Proceedings of the National Academy of Sciences of the United States of America, 108(5), 1862–1867. [CrossRef]
  10. Bertram, K., Agafonov, D. E., Dybkov, O., Haselbach, D., Leelaram, M. N., Will, C. L., Urlaub, H., Kastner, B., Lührmann, R., & Stark, H. (2017b). Cryo-EM Structure of a Pre-catalytic Human Spliceosome Primed for Activation. Cell, 170(4), 701–713.e11. [CrossRef]
  11. Bertram, K., Agafonov, D. E., Liu, W. T., Dybkov, O., Will, C. L., Hartmuth, K., Urlaub, H., Kastner, B., Stark, H., & Lührmann, R. (2017a). Cryo-EM structure of a human spliceosome activated for step 2 of splicing. Nature, 542(7641), 318–323. [CrossRef]
  12. Biancon, G., Joshi, P., Zimmer, J. T., Hunck, T., Gao, Y., Lessard, M. D., Courchaine, E., Barentine, A. E. S., Machyna, M., Botti, V., Qin, A., Gbyli, R., Patel, A., Song, Y., Kiefer, L., Viero, G., Neuenkirchen, N., Lin, H., Bewersdorf, J., Simon, M. D., … Halene, S. (2022). Precision analysis of mutant U2AF1 activity reveals deployment of stress granules in myeloid malignancies. Molecular cell, 82(6), 1107–1122.e7. [CrossRef]
  13. Boudvillain, M., de Lencastre, A., & Pyle, A. M. (2000). A tertiary interaction that links active-site domains to the 5' splice site of a group II intron. Nature, 406(6793), 315–318. [CrossRef]
  14. Bousquets-Muñoz, P., Díaz-Navarro, A., Nadeu, F., Sánchez-Pitiot, A., López-Tamargo, S., Shuai, S., Balbín, M., Tubio, J. M. C., Beà, S., Martin-Subero, J. I., Gutiérrez-Fernández, A., Stein, L. D., Campo, E., & Puente, X. S. (2022). PanCancer analysis of somatic mutations in repetitive regions reveals recurrent mutations in snRNA U2. NPJ genomic medicine, 7(1), 19. [CrossRef]
  15. Bradley, R. K., Merkin, J., Lambert, N. J., & Burge, C. B. (2012). Alternative splicing of RNA triplets is often regulated and accelerates proteome evolution. PLoS biology, 10(1), e1001229. [CrossRef]
  16. Briese, M., Haberman, N., Sibley, C. R., Faraway, R., Elser, A. S., Chakrabarti, A. M., Wang, Z., König, J., Perera, D., Wickramasinghe, V. O., Venkitaraman, A. R., Luscombe, N. M., Saieva, L., Pellizzoni, L., Smith, C. W. J., Curk, T., & Ule, J. (2019). A systems view of spliceosomal assembly and branchpoints with iCLIP. Nature structural & molecular biology, 26(10), 930–940. [CrossRef]
  17. Buerer, L., Clark, N. E., Welch, A., Duan, C., Taggart, A. J., Townley, B. A., Wang, J., Soemedi, R., Rong, S., Lin, C. L., Zeng, Y., Katolik, A., Staley, J. P., Damha, M. J., Mosammaparast, N., & Fairbrother, W. G. (2024). The debranching enzyme Dbr1 regulates lariat turnover and intron splicing. Nature communications, 15(1), 4617. [CrossRef]
  18. Carmel, I., Tal, S., Vig, I., & Ast, G. (2004). Comparative analysis detects dependencies among the 5' splice-site positions. RNA (New York, N.Y.), 10(5), 828–840. [CrossRef]
  19. Chanfreau, G., & Jacquier, A. (1993). Interaction of intronic boundaries is required for the second splicing step efficiency of a group II intron. The EMBO journal, 12(13), 5173–5180. [CrossRef]
  20. Chanfreau, G., Legrain, P., Dujon, B., & Jacquier, A. (1994). Interaction between the first and last nucleotides of pre-mRNA introns is a determinant of 3' splice site selection in S. cerevisiae. Nucleic acids research, 22(11), 1981–1987. [CrossRef]
  21. Charenton, C., Wilkinson, M. E., & Nagai, K. (2019). Mechanism of 5' splice site transfer for human spliceosome activation. Science (New York, N.Y.), 364(6438), 362–367. [CrossRef]
  22. Charette, M., & Gray, M. W. (2000). Pseudouridine in RNA: what, where, how, and why. IUBMB life, 49(5), 341–351. [CrossRef]
  23. Chen, C., Zhao, X., Kierzek, R., & Yu, Y. T. (2010). A flexible RNA backbone within the polypyrimidine tract is required for U2AF65 binding and pre-mRNA splicing in vivo. Molecular and cellular biology, 30(17), 4108–4119. [CrossRef]
  24. Chen, J. Y., Stands, L., Staley, J. P., Jackups, R. R., Jr, Latus, L. J., & Chang, T. H. (2001). Specific alterations of U1-C protein or U1 small nuclear RNA can eliminate the requirement of Prp28p, an essential DEAD box splicing factor. Molecular cell, 7(1), 227–232. [CrossRef]
  25. Chen, Y., Dawes, R., Kim, H. C., Ljungdahl, A., Stenton, S. L., Walker, S., Lord, J., Lemire, G., Martin-Geary, A. C., Ganesh, V. S., Ma, J., Ellingford, J. M., Delage, E., D'Souza, E. N., Dong, S., Adams, D. R., Allan, K., Bakshi, M., Baldwin, E. E., Berger, S. I., … Whiffin, N. (2024). De novo variants in the RNU4-2 snRNA cause a frequent neurodevelopmental syndrome. Nature, 632(8026), 832–840. [CrossRef]
  26. Chiu, A. C., Suzuki, H. I., Wu, X., Mahat, D. B., Kriz, A. J., & Sharp, P. A. (2018). Transcriptional Pause Sites Delineate Stable Nucleosome-Associated Premature Polyadenylation Suppressed by U1 snRNP. Molecular cell, 69(4), 648–663.e7. [CrossRef]
  27. Cohen, J. B., Snow, J. E., Spencer, S. D., & Levinson, A. D. (1994). Suppression of mammalian 5' splice-site defects by U1 small nuclear RNAs from a distance. Proceedings of the National Academy of Sciences of the United States of America, 91(22), 10470–10474. [CrossRef]
  28. Corrionero, A., Raker, V. A., Izquierdo, J. M., & Valcárcel, J. (2011). Strict 3' splice site sequence requirements for U2 snRNP recruitment after U2AF binding underlie a genetic defect leading to autoimmune disease. RNA (New York, N.Y.), 17(3), 401–411. [CrossRef]
  29. Cortes, J. J., Sontheimer, E. J., Seiwert, S. D., & Steitz, J. A. (1993). Mutations in the conserved loop of human U5 snRNA generate use of novel cryptic 5' splice sites in vivo. The EMBO journal, 12(13), 5181–5189. [CrossRef]
  30. Costa M. (2022). Group II Introns: Flexibility and Repurposing. Frontiers in molecular biosciences, 9, 916157. [CrossRef]
  31. Costa, M., Walbott, H., Monachello, D., Westhof, E., & Michel, F. (2016). Crystal structures of a group II intron lariat primed for reverse splicing. Science (New York, N.Y.), 354(6316), aaf9258. [CrossRef]
  32. Covello, G., Ibrahim, G. H., Bacchi, N., Casarosa, S., & Denti, M. A. (2022). Exon Skipping Through Chimeric Antisense U1 snRNAs to Correct Retinitis Pigmentosa GTPase-Regulator (RPGR) Splice Defect. Nucleic acid therapeutics, 32(4), 333–349. [CrossRef]
  33. Cretu, C., Gee, P., Liu, X., Agrawal, A., Nguyen, T. V., Ghosh, A. K., Cook, A., Jurica, M., Larsen, N. A., & Pena, V. (2021). Structural basis of intron selection by U2 snRNP in the presence of covalent inhibitors. Nature communications, 12(1), 4491. [CrossRef]
  34. Crispino, J. D., & Sharp, P. A. (1995). A U6 snRNA:pre-mRNA interaction can be rate-limiting for U1-independent splicing. Genes & development, 9(18), 2314–2323. [CrossRef]
  35. Damianov, A., Lin, C. H., Zhang, J., Manley, J. L., & Black, D. L. (2025). Cancer-associated SF3B1 mutation K700E causes widespread changes in U2/branchpoint recognition without altering splicing. Proceedings of the National Academy of Sciences of the United States of America, 122(13), e2423776122. [CrossRef]
  36. Deirdre, A., Scadden, J., & Smith, C. W. (1995). Interactions between the terminal bases of mammalian introns are retained in inosine-containing pre-mRNAs. The EMBO journal, 14(13), 3236–3246. [CrossRef]
  37. Depienne, C., & Mandel, J. L. (2021). 30 years of repeat expansion disorders: What have we learned and what are the remaining challenges?. American journal of human genetics, 108(5), 764–785. [CrossRef]
  38. Donadon, I., Bussani, E., Riccardi, F., Licastro, D., Romano, G., Pianigiani, G., Pinotti, M., Konstantinova, P., Evers, M., Lin, S., Rüegg, M. A., & Pagani, F. (2019). Rescue of spinal muscular atrophy mouse models with AAV9-Exon-specific U1 snRNA. Nucleic acids research, 47(14), 7618–7632. [CrossRef]
  39. Dong, X., Ranganathan, S., Qu, G., Piazza, C. L., & Belfort, M. (2018). Structural accommodations accompanying splicing of a group II intron RNP. Nucleic acids research, 46(16), 8542–8556. [CrossRef]
  40. Fabrizio, P., & Abelson, J. (1990). Two domains of yeast U6 small nuclear RNA required for both steps of nuclear precursor messenger RNA splicing. Science (New York, N.Y.), 250(4979), 404–409. [CrossRef]
  41. Fernandez Alanis, E., Pinotti, M., Dal Mas, A., Balestra, D., Cavallari, N., Rogalska, M. E., Bernardi, F., & Pagani, F. (2012). An exon-specific U1 small nuclear RNA (snRNA) strategy to correct splicing defects. Human molecular genetics, 21(11), 2389–2398. [CrossRef]
  42. Fica, S. M., Oubridge, C., Galej, W. P., Wilkinson, M. E., Bai, X. C., Newman, A. J., & Nagai, K. (2017). Structure of a spliceosome remodelled for exon ligation. Nature, 542(7641), 377–380. [CrossRef]
  43. Fica, S. M., Oubridge, C., Wilkinson, M. E., Newman, A. J., & Nagai, K. (2019). A human postcatalytic spliceosome structure reveals essential roles of metazoan factors for exon ligation. Science (New York, N.Y.), 363(6428), 710–714. [CrossRef]
  44. Frison-Roche, C., Demier, C. M., Cottin, S., Lainé, J., Arandel, L., Halliez, M., Lemaitre, M., Lornage, X., Strochlic, L., Swanson, M. S., Martinat, C., Messéant, J., Furling, D., & Rau, F. (2025). MBNL deficiency in motor neurons disrupts neuromuscular junction maintenance and gait coordination. Brain : a journal of neurology, 148(4), 1180–1193. [CrossRef]
  45. Fu, Y., Masuda, A., Ito, M., Shinmi, J., & Ohno, K. (2011). AG-dependent 3'-splice sites are predisposed to aberrant splicing due to a mutation at the first nucleotide of an exon. Nucleic acids research, 39(10), 4396–4404. [CrossRef]
  46. Fukumura, K., & Inoue, K. (2009). Role and mechanism of U1-independent pre-mRNA splicing in the regulation of alternative splicing. RNA biology, 6(4), 395–398. [CrossRef]
  47. Fukumura, K., Taniguchi, I., Sakamoto, H., Ohno, M., & Inoue, K. (2009). U1-independent pre-mRNA splicing contributes to the regulation of alternative splicing. Nucleic acids research, 37(6), 1907–1914. [CrossRef]
  48. Galej, W. P., Nguyen, T. H., Newman, A. J., & Nagai, K. (2014). Structural studies of the spliceosome: zooming into the heart of the machine. Current opinion in structural biology, 25(100), 57–66. [CrossRef]
  49. Galej, W. P., Wilkinson, M. E., Fica, S. M., Oubridge, C., Newman, A. J., & Nagai, K. (2016). Cryo-EM structure of the spliceosome immediately after branching. Nature, 537(7619), 197–201. [CrossRef]
  50. Glasser, E., Maji, D., Biancon, G., Puthenpeedikakkal, A. M. K., Cavender, C. E., Tebaldi, T., Jenkins, J. L., Mathews, D. H., Halene, S., & Kielkopf, C. L. (2022). Pre-mRNA splicing factor U2AF2 recognizes distinct conformations of nucleotide variants at the center of the pre-mRNA splice site signal. Nucleic acids research, 50(9), 5299–5312. [CrossRef]
  51. Gooding, C., Clark, F., Wollerton, M. C., Grellscheid, S. N., Groom, H., & Smith, C. W. (2006). A class of human exons with predicted distant branch points revealed by analysis of AG dinucleotide exclusion zones. Genome biology, 7(1), R1. [CrossRef]
  52. Grainger, R. J., & Beggs, J. D. (2005). Prp8 protein: at the heart of the spliceosome. RNA (New York, N.Y.), 11(5), 533–557. [CrossRef]
  53. Greene, D., Thys, C., Berry, I. R., Jarvis, J., Ortibus, E., Mumford, A. D., Freson, K., & Turro, E. (2024). Mutations in the U4 snRNA gene RNU4-2 cause one of the most prevalent monogenic neurodevelopmental disorders. Nature medicine, 30(8), 2165–2169. [CrossRef]
  54. Gregorich, Z. R., Zhang, Y., Kamp, T. J., Granzier, H. L., & Guo, W. (2024). Mechanisms of RBM20 Cardiomyopathy: Insights From Model Systems. Circulation. Genomic and precision medicine, 17(1), e004355. [CrossRef]
  55. Hamid, F. M., & Makeyev, E. V. (2017). A mechanism underlying position-specific regulation of alternative splicing. Nucleic acids research, 45(21), 12455–12468. [CrossRef]
  56. Haselbach, D., Komarov, I., Agafonov, D. E., Hartmuth, K., Graf, B., Dybkov, O., Urlaub, H., Kastner, B., Lührmann, R., & Stark, H. (2018). Structure and Conformational Dynamics of the Human Spliceosomal Bact Complex. Cell, 172(3), 454–464.e11. [CrossRef]
  57. Hatch, S. T., Smargon, A. A., & Yeo, G. W. (2022). Engineered U1 snRNAs to modulate alternatively spliced exons. Methods (San Diego, Calif.), 205, 140–148. [CrossRef]
  58. Hodson, M. J., Hudson, A. J., Cherny, D., & Eperon, I. C. (2012). The transition in spliceosome assembly from complex E to complex A purges surplus U1 snRNPs from alternative splice sites. Nucleic acids research, 40(14), 6850–6862. https://www.nakb.org/ndbmodule/bp-catalog/ [Accessed July 23, 2026]. [CrossRef]
  59. Hwang, D. Y., & Cohen, J. B. (1996). U1 snRNA promotes the selection of nearby 5' splice sites by U6 snRNA in mammalian cells. Genes & development, 10(3), 338–350. [CrossRef]
  60. Ichiyanagi, K., Beauregard, A., Lawrence, S., Smith, D., Cousineau, B., & Belfort, M. (2002). Retrotransposition of the Ll.LtrB group II intron proceeds predominantly via reverse splicing into DNA targets. Molecular microbiology, 46(5), 1259–1272. [CrossRef]
  61. Irimia, M., & Roy, S. W. (2008). Evolutionary convergence on highly-conserved 3' intron structures in intron-poor eukaryotes and insights into the ancestral eukaryotic genome. PLoS genetics, 4(8), e1000148. [CrossRef]
  62. Ishigami, Y., Ohira, T., Isokawa, Y., Suzuki, Y., & Suzuki, T. (2021). A single m6A modification in U6 snRNA diversifies exon sequence at the 5' splice site. Nature communications, 12(1), 3244. [CrossRef]
  63. Jackson, A., Thaker, N., Blakes, A., Rice, G., Griffiths-Jones, S., Balasubramanian, M., Campbell, J., Shannon, N., Choi, J., Hong, J., Hunt, D., de Burca, A., Kim, S. Y., Kim, T., Lee, S., Redman, M., Rius, R., Simons, C., Tan, T. Y., Ellingford, J., … Banka, S. (2025). Analysis of R-loop forming regions identifies RNU2-2 and RNU5B-1 as neurodevelopmental disorder genes. Nature genetics, 57(6), 1362–1366. [CrossRef]
  64. Jacquier, A., & Michel, F. (1990). Base-pairing interactions involving the 5' and 3'-terminal nucleotides of group II self-splicing introns. Journal of molecular biology, 213(3), 437–447. [CrossRef]
  65. Jakab, G., Mougin, A., Kis, M., Pollák, T., Antal, M., Branlant, C., & Solymosy, F. (1997). Chlamydomonas U2, U4 and U6 snRNAs. An evolutionary conserved putative third interaction between U4 and U6 snRNAs which has a counterpart in the U4atac-U6atac snRNA duplex. Biochimie, 79(7), 387–395. [CrossRef]
  66. Jobbins, A. M., Campagne, S., Weinmeister, R., Lucas, C. M., Gosliga, A. R., Clery, A., Chen, L., Eperon, L. P., Hodson, M. J., Hudson, A. J., Allain, F. H. T., & Eperon, I. C. (2022). Exon-independent recruitment of SRSF1 is mediated by U1 snRNP stem-loop 3. The EMBO journal, 41(1), e107640. [CrossRef]
  67. Jüschke, C., Klopstock, T., Catarino, C. B., Owczarek-Lipska, M., Wissinger, B., & Neidhardt, J. (2021). Autosomal dominant optic atrophy: A novel treatment for OPA1 splice defects using U1 snRNA adaption. Molecular therapy. Nucleic acids, 26, 1186–1197. [CrossRef]
  68. Keating, K. S., Toor, N., Perlman, P. S., & Pyle, A. M. (2010). A structural analysis of the group II intron active site and implications for the spliceosome. RNA (New York, N.Y.), 16(1), 1–9. [CrossRef]
  69. Kent, O. A., Reayi, A., Foong, L., Chilibeck, K. A., & MacMillan, A. M. (2003). Structuring of the 3' splice site by U2AF65. The Journal of biological chemistry, 278(50), 50572–50577. [CrossRef]
  70. Kim, C. H., & Abelson, J. (1996). Site-specific crosslinks of yeast U6 snRNA to the pre-mRNA near the 5' splice site. RNA (New York, N.Y.), 2(10), 995–1010.
  71. Kirchhoff, C. L., Powell, H. R., Galardi, J. W., Loerch, S., Pulvino, M. J., Jenkins, J. L., Boutz, P. L., Kielkopf, C. L. (2026). SF3B1 phosphorylation prompts U2AF2 dissociation for widespread control of pre-mRNA splicing. bioRxiv 2026.01.19.700466; [CrossRef]
  72. Konarska, M. M., Vilardell, J., & Query, C. C. (2006). Repositioning of the reaction intermediate within the catalytic center of the spliceosome. Molecular cell, 21(4), 543–553. [CrossRef]
  73. Lee, S. A., Ogawa, M., Saito, Y., Shimazaki, R., Awaya, T., Hosokawa, M., Kurosawa, R., Ohara, H., Takeuchi, A., Hayashi, S., Goto, Y. I., Hagiwara, M., Nishino, I., & Noguchi, S. (2025). Substitutions of nucleotides at the 3' ends of COL6A1/2/3 exons induce exon skipping associated with collagen VI-related muscular dystrophies and therapeutic strategies. Genetics in medicine : official journal of the American College of Medical Genetics, 27(7), 101431. [CrossRef]
  74. Lee, S., & Stevens, S. W. (2016). Spliceosomal intronogenesis. Proceedings of the National Academy of Sciences of the United States of America, 113(23), 6514–6519. [CrossRef]
  75. Leontis, N. B., Stombaugh, J., & Westhof, E. (2002). The non-Watson-Crick base pairs and their associated isostericity matrices. Nucleic acids research, 30(16), 3497–3531. [CrossRef]
  76. Liu, N., Dong, X., Hu, C., Zeng, J., Wang, J., Wang, J., Wang, H. W., & Belfort, M. (2020). Exon and protein positioning in a pre-catalytic group II intron RNP primed for splicing. Nucleic acids research, 48(19), 11185–11198. [CrossRef]
  77. Lund, M., & Kjems, J. (2002). Defining a 5' splice site by functional selection in the presence and absence of U1 snRNA 5' end. RNA (New York, N.Y.), 8(2), 166–179. [CrossRef]
  78. Martelly, W., Fellows, B., Kang, P., Vashisht, A., Wohlschlegel, J. A., & Sharma, S. (2021). Synergistic roles for human U1 snRNA stem-loops in pre-mRNA splicing. RNA biology, 18(12), 2576–2593. [CrossRef]
  79. Masquida, B., & Westhof, E. (2000). On the wobble GoU and related pairs. RNA (New York, N.Y.), 6(1), 9–15. [CrossRef]
  80. Mercer, T. R., Clark, M. B., Andersen, S. B., Brunck, M. E., Haerty, W., Crawford, J., Taft, R. J., Nielsen, L. K., Dinger, M. E., & Mattick, J. S. (2015). Genome-wide discovery of human splicing branchpoints. Genome research, 25(2), 290–303. [CrossRef]
  81. Michel, F., & Ferat, J. L. (1995). Structure and activities of group II introns. Annual review of biochemistry, 64, 435–461. [CrossRef]
  82. Michel, F., Umesono, K., & Ozeki, H. (1989). Comparative and functional anatomy of group II catalytic introns--a review. Gene, 82(1), 5–30. [CrossRef]
  83. Mimoso, C. A., & Adelman, K. (2023). U1 snRNP increases RNA Pol II elongation rate to enable synthesis of long genes. Molecular cell, 83(8), 1264–1279.e10. [CrossRef]
  84. Moafinejad, S. N., de Aquino, B. R. H., Boniecki, M. J., Pandaranadar Jeyeram, I. P. N., Nikolaev, G., Magnus, M., Farsani, M. A., Badepally, N. G., Wirecki, T. K., Stefaniak, F., & Bujnicki, J. M. (2024). SimRNAweb v2.0: a web server for RNA folding simulations and 3D structure modeling, with optional restraints and enhanced analysis of folding trajectories. Nucleic acids research, 52(W1), W368–W373. [CrossRef]
  85. Montañés-Agudo, P., Aufiero, S., Schepers, E. N., van der Made, I., Cócera-Ortega, L., Ernault, A. C., Richard, S., Kuster, D. W. D., Christoffels, V. M., Pinto, Y. M., & Creemers, E. E. (2023). The RNA-binding protein QKI governs a muscle-specific alternative splicing program that shapes the contractile function of cardiomyocytes. Cardiovascular research, 119(5), 1161–1174. [CrossRef]
  86. NAKB Nucleic Acid Knowledgebase. (2026) RNA Basepair Catalog.
  87. Nameki, N., Takizawa, M., Suzuki, T., Tani, S., Kobayashi, N., Sakamoto, T., Muto, Y., & Kuwasako, K. (2022). Structural basis for the interaction between the first SURP domain of the SF3A1 subunit in U2 snRNP and the human splicing factor SF1. Protein science : a publication of the Protein Society, 31(10), e4437. [CrossRef]
  88. Nava, C., Cogne, B., Santini, A., Leitão, E., Lecoquierre, F., Chen, Y., Stenton, S. L., Besnard, T., Heide, S., Baer, S., Jakhar, A., Neuser, S., Keren, B., Faudet, A., Forlani, S., Faoucher, M., Uguen, K., Platzer, K., Afenjar, A., Alessandri, J. L., … Depienne, C. (2025). Dominant variants in major spliceosome U4 and U5 small nuclear RNA genes cause neurodevelopmental disorders through splicing disruption. Nature genetics, 57(6), 1374–1388. [CrossRef]
  89. Negi, M. S., Krishnan, V. P., Saraf, N., & Vijayraghavan, U. (2025). Prp16 enables efficient splicing of introns with diverse exonic consensus elements in the short-intron rich Cryptococcus neoformans transcriptome. RNA biology, 22(1), 1–14. [CrossRef]
  90. Newman, A. J., Teigelkamp, S., & Beggs, J. D. (1995). snRNA interactions at 5' and 3' splice sites monitored by photoactivated crosslinking in yeast spliceosomes. RNA (New York, N.Y.), 1(9), 968–980. https://rnajournal.cshlp.org/content/1/9/968.full.pdf.
  91. Nielsen, K. H., & Staley, J. P. (2012). Spliceosome activation: U4 is the path, stem I is the goal, and Prp8 is the keeper. Let's cheer for the ATPase Brr2!. Genes & development, 26(22), 2461–2467. [CrossRef]
  92. Nissen, P., Ippolito, J. A., Ban, N., Moore, P. B., & Steitz, T. A. (2001). RNA tertiary interactions in the large ribosomal subunit: the A-minor motif. Proceedings of the National Academy of Sciences of the United States of America, 98(9), 4899–4903. [CrossRef]
  93. Obuća, M., Cvačková, Z., Kubovčiak, J., Kolář, M., & Staněk, D. (2022). Retinitis pigmentosa-linked mutation in DHX38 modulates its splicing activity. PloS one, 17(4), e0265742. [CrossRef]
  94. Ohno, K., Tsujino, A., Shen, X. M., Milone, M., & Engel, A. G. (2005). Spectrum of splicing errors caused by CHRNE mutations affecting introns and intron/exon boundaries. Journal of medical genetics, 42(8), e53. [CrossRef]
  95. Olthof, A. M., Schwoerer, C. F., Girardini, K. N., Weber, A. L., Doggett, K., Mieruszynski, S., Heath, J. K., Moore, T. E., Biran, J., & Kanadia, R. N. (2024). Taxonomy of introns and the evolution of minor introns. Nucleic acids research, 52(15), 9247–9266. [CrossRef]
  96. Paggi, J. M., & Bejerano, G. (2018). A sequence-based, deep learning model accurately predicts RNA splicing branchpoints. RNA (New York, N.Y.), 24(12), 1647–1658. [CrossRef]
  97. Parker, M. T., Fica, S. M., & Simpson, G. G. (2025). RNA splicing: a split consensus reveals two major 5' splice site classes. Open biology, 15(1), 240293. [CrossRef]
  98. Parker, M. T., Soanes, B. K., Kusakina, J., Larrieu, A., Knop, K., Joy, N., Breidenbach, F., Sherwood, A. V., Barton, G. J., Fica, S. M., Davies, B. H., & Simpson, G. G. (2022). m6A modification of U6 snRNA modulates usage of two major classes of pre-mRNA 5' splice site. eLife, 11, e78808. [CrossRef]
  99. Parker, R., & Siliciano, P. G. (1993). Evidence for an essential non-Watson-Crick interaction between the first and last nucleotides of a nuclear pre-mRNA intron. Nature, 361(6413), 660–662. [CrossRef]
  100. Peled-Zehavi, H., Berglund, J. A., Rosbash, M., & Frankel, A. D. (2001). Recognition of RNA branch point sequences by the KH domain of splicing factor 1 (mammalian branch point binding protein) in a splicing factor complex. Molecular and cellular biology, 21(15), 5232–5241. [CrossRef]
  101. Pendleton, K. E., Chen, B., Liu, K., Hunter, O. V., Xie, Y., Tu, B. P., & Conrad, N. K. (2017). The U6 snRNA m6A Methyltransferase METTL16 Regulates SAM Synthetase Intron Retention. Cell, 169(5), 824–835.e14. [CrossRef]
  102. Pereira de Castro, K. L., Abril, J. M., Liao, K. C., Hao, H., Donohue, J. P., Russell, W. K., & Fagg, W. S. (2024). Competition for the conserved branch point sequence influences physiological outcomes in pre-mRNA splicing eLife 13:RP103167. [CrossRef]
  103. Perriman, R., & Ares, M., Jr (2010). Invariant U2 snRNA nucleotides form a stem loop to recognize the intron early in splicing. Molecular cell, 38(3), 416–427. [CrossRef]
  104. Rauhut, R., Fabrizio, P., Dybkov, O., Hartmuth, K., Pena, V., Chari, A., Kumar, V., Lee, C. T., Urlaub, H., Kastner, B., Stark, H., & Lührmann, R. (2016). Molecular architecture of the Saccharomyces cerevisiae activated spliceosome. Science (New York, N.Y.), 353(6306), 1399–1405. [CrossRef]
  105. Riley, L. A., Zhang, X., Douglas, C. M., Mijares, J. M., Hammers, D. W., Wolff, C. A., Wood, N. B., Olafson, H. R., Du, P., Labeit, S., Previs, M. J., Wang, E. T., & Esser, K. A. (2022). The skeletal muscle circadian clock regulates titin splicing through RBM20. eLife, 11, e76478. [CrossRef]
  106. Romano, G., Riccardi, F., Bussani, E., Vodret, S., Licastro, D., Ragone, I., Ronzitti, G., Morini, E., Slaugenhaupt, S. A., & Pagani, F. (2022). Rescue of a familial dysautonomia mouse model by AAV9-Exon-specific U1 snRNA. American journal of human genetics, 109(8), 1534–1548. [CrossRef]
  107. Rozov, A., Demeshkina, N., Westhof, E., Yusupov, M., & Yusupova, G. (2015). Structural insights into the translational infidelity mechanism. Nature communications, 6, 7251. [CrossRef]
  108. Rypniewski, W., Banaszak, K., Kuliński, T., & Kiliszek, A. (2016). Watson-Crick-like pairs in CCUG repeats: evidence for tautomeric shifts or protonation. RNA (New York, N.Y.), 22(1), 22–31. [CrossRef]
  109. Sakers, K., Liu, Y., Llaci, L., Lee, S. M., Vasek, M. J., Rieger, M. A., Brophy, S., Tycksen, E., Lewis, R., Maloney, S. E., & Dougherty, J. D. (2021). Loss of Quaking RNA binding protein disrupts the expression of genes associated with astrocyte maturation in mouse brain. Nature communications, 12(1), 1537. [CrossRef]
  110. Savarese, M., Sarparanta, J., Vihola, A., Udd, B., & Hackman, P. (2016). Increasing Role of Titin Mutations in Neuromuscular Disorders. Journal of neuromuscular diseases, 3(3), 293–308. [CrossRef]
  111. Ségault, V., Will, C. L., Polycarpou-Schwarz, M., Mattaj, I. W., Branlant, C., & Lührmann, R. (1999). Conserved loop I of U5 small nuclear RNA is dispensable for both catalytic steps of pre-mRNA splicing in HeLa nuclear extracts. Molecular and cellular biology, 19(4), 2782–2790. [CrossRef]
  112. Sellier, C., Cerro-Herreros, E., Blatter, M., Freyermuth, F., Gaucherot, A., Ruffenach, F., Sarkar, P., Puymirat, J., Udd, B., Day, J. W., Meola, G., Bassez, G., Fujimura, H., Takahashi, M. P., Schoser, B., Furling, D., Artero, R., Allain, F. H. T., Llamusi, B., & Charlet-Berguerand, N. (2018). rbFOX1/MBNL1 competition for CCUG RNA repeats binding contributes to myotonic dystrophy type 1/type 2 differences. Nature communications, 9(1), 2009. [CrossRef]
  113. Sickmier, E. A., Frato, K. E., Shen, H., Paranawithana, S. R., Green, M. R., & Kielkopf, C. L. (2006). Structural basis for polypyrimidine tract recognition by the essential pre-mRNA splicing factor U2AF65. Molecular cell, 23(1), 49–59. [CrossRef]
  114. Slat, V. A. Exploration of the Cyanidioschyzon merolae intron landscape. (2021) MSc Thesis, UNBC https://dam-oclc.bac-lac.gc.ca/download?is_thesis=1&oclc_number=1287010185&id=89a0257e-ccc6-4268-bcae-c195415ba3b8&fileName=file.pdf.
  115. Smith, D. J., & Konarska, M. M. (2008). Mechanistic insights from reversible splicing catalysis. RNA (New York, N.Y.), 14(10), 1975–1978. [CrossRef]
  116. Sontheimer, E. J., & Steitz, J. A. (1993). The U5 and U6 small nuclear RNAs as active site components of the spliceosome. Science (New York, N.Y.), 262(5142), 1989–1996. [CrossRef]
  117. Stark, M. R., Dunn, E. A., Dunn, W. S., Grisdale, C. J., Daniele, A. R., Halstead, M. R., Fast, N. M., & Rader, S. D. (2015). Dramatically reduced spliceosome in Cyanidioschyzon merolae. Proceedings of the National Academy of Sciences of the United States of America, 112(11), E1191–E1200. [CrossRef]
  118. Šulc, P., Ouldridge, T. E., Romano, F., Doye, J. P., & Louis, A. A. (2015). Modelling toehold-mediated RNA strand displacement. Biophysical journal, 108(5), 1238–1247. [CrossRef]
  119. Tourasse, N. J., Stabell, F. B., & Kolstø, A. B. (2011). Diversity, mobility, and structural and functional evolution of group II introns carrying an unusual 3' extension. BMC research notes, 4, 564. [CrossRef]
  120. Tseng, C. K., & Cheng, S. C. (2008). Both catalytic steps of nuclear pre-mRNA splicing are reversible. Science (New York, N.Y.), 320(5884), 1782–1784. [CrossRef]
  121. Ule, J., & Blencowe, B. J. (2019). Alternative Splicing Regulatory Networks: Functions, Mechanisms, and Evolution. Molecular cell, 76(2), 329–345. [CrossRef]
  122. Wan, R., Bai, R., Yan, C., Lei, J., & Shi, Y. (2019). Structures of the Catalytically Activated Yeast Spliceosome Reveal the Mechanism of Branching. Cell, 177(2), 339–351.e13. [CrossRef]
  123. Wan, R., Yan, C., Bai, R., Huang, G., & Shi, Y. (2016b). Structure of a yeast catalytic step I spliceosome at 3.4 Å resolution. Science (New York, N.Y.), 353(6302), 895–904. [CrossRef]
  124. Wan, R., Yan, C., Bai, R., Lei, J., & Shi, Y. (2017). Structure of an Intron Lariat Spliceosome from Saccharomyces cerevisiae. Cell, 171(1), 120–132.e12. [CrossRef]
  125. Wan, R., Yan, C., Bai, R., Wang, L., Huang, M., Wong, C. C., & Shi, Y. (2016a). The 3.8 Å structure of the U4/U6.U5 tri-snRNP: Insights into spliceosome assembly and catalysis. Science (New York, N.Y.), 351(6272), 466–475. [CrossRef]
  126. Wang, W., Hellinga, H. W., & Beese, L. S. (2011). Structural evidence for the rare tautomer hypothesis of spontaneous mutagenesis. Proceedings of the National Academy of Sciences of the United States of America, 108(43), 17644–17648. [CrossRef]
  127. Wilkinson, M. E., Charenton, C., & Nagai, K. (2020). RNA Splicing by the Spliceosome. Annual review of biochemistry, 89, 359–388. [CrossRef]
  128. Wilkinson, M. E., Fica, S. M., Galej, W. P., & Nagai, K. (2021). Structural basis for conformational equilibrium of the catalytic spliceosome. Molecular cell, 81(7), 1439–1452.e9. [CrossRef]
  129. Wilkinson, M. E., Fica, S. M., Galej, W. P., Norman, C. M., Newman, A. J., & Nagai, K. (2017). Postcatalytic spliceosome structure reveals mechanism of 3'-splice site selection. Science (New York, N.Y.), 358(6368), 1283–1288. [CrossRef]
  130. Xie, J., Wang, L., & Lin, R. J. (2023). Variations of intronic branchpoint motif: identification and functional implications in splicing and disease. Communications biology, 6(1), 1142. [CrossRef]
  131. Xu, L., Liu, T., Chung, K., & Pyle, A. M. (2023). Structural insights into intron catalysis and dynamics during splicing. Nature, 624(7992), 682–688. [CrossRef]
  132. Yan, C., Wan, R., Bai, R., Huang, G., & Shi, Y. (2016). Structure of a yeast activated spliceosome at 3.5 Å resolution. Science (New York, N.Y.), 353(6302), 904–911. [CrossRef]
  133. Yan, C., Wan, R., Bai, R., Huang, G., & Shi, Y. (2017). Structure of a yeast step II catalytically activated spliceosome. Science (New York, N.Y.), 355(6321), 149–155. [CrossRef]
  134. Yang, H., Beutler, B., & Zhang, D. (2022). Emerging roles of spliceosome in cancer and immunity. Protein & cell, 13(8), 559–579. [CrossRef]
  135. Yoshida, H., Park, S. Y., Sakashita, G., Nariai, Y., Kuwasako, K., Muto, Y., Urano, T., & Obayashi, E. (2020). Elucidation of the aberrant 3' splice site selection by cancer-associated mutations on the U2AF1. Nature communications, 11(1), 4744. [CrossRef]
  136. Zhan, X., Yan, C., Zhang, X., Lei, J., & Shi, Y. (2018b). Structures of the human pre-catalytic spliceosome and its precursor spliceosome. Cell research, 28(12), 1129–1140. [CrossRef]
  137. Zhan, X., Yan, C., Zhang, X., Lei, J., & Shi, Y. (2018a). Structure of a human catalytic step I spliceosome. Science (New York, N.Y.), 359(6375), 537–545. [CrossRef]
  138. Zhang, W., Zhang, X., Zhan, X., Bai, R., Lei, J., Yan, C., & Shi, Y. (2024). Structural insights into human exon-defined spliceosome prior to activation. Cell research, 34(6), 428–439. [CrossRef]
  139. Zhang, X., Yan, C., Hang, J., Finci, L. I., Lei, J., & Shi, Y. (2017). An Atomic Structure of the Human Spliceosome. Cell, 169(5), 918–929.e14. [CrossRef]
  140. Zhang, X., Yan, C., Zhan, X., Li, L., Lei, J., & Shi, Y. (2018). Structure of the human activated spliceosome in three conformational states. Cell research, 28(3), 307–322. [CrossRef]
  141. Zhang, X., Zhan, X., Yan, C., Zhang, W., Liu, D., Lei, J., & Shi, Y. (2019). Structures of the human spliceosomes before and after release of the ligated exon. Cell research, 29(4), 274–285. [CrossRef]
  142. Zhang, Y., Yin, C., Wang, Y., Yan, K., Zhao, B., Hu, X., Wan, Y., Cheng, H., & Huang, J. (2025). A common structural mechanism for RNA recognition by the SF3B complex in mRNA splicing and export. Nucleic acids research, 53(15), gkaf759. [CrossRef]
  143. Zhang, Z., Kumar, V., Dybkov, O., Will, C. L., Urlaub, H., Stark, H., & Lührmann, R. (2024a). Cryo-EM analyses of dimerized spliceosomes provide new insights into the functions of B complex proteins. The EMBO journal, 43(6), 1065–1088. [CrossRef]
  144. Zhang, Z., Kumar, V., Dybkov, O., Will, C. L., Zhong, J., Ludwig, S. E. J., Urlaub, H., Kastner, B., Stark, H., & Lührmann, R. (2024b). Structural insights into the cross-exon to cross-intron spliceosome switch. Nature, 630(8018), 1012–1019. [CrossRef]
  145. Zhuang, Y., & Weiner, A. M. (1986). A compensatory base change in U1 snRNA suppresses a 5' splice site mutation. Cell, 46(6), 827–835. [CrossRef]
Figure 7. Structural model of the U5 Loop1 base-pairing with the exons and the non-Watson-Crick pair between the ends of the intron. The model was generated by SimRNA v2.0 (Moafinejad et al., 2024) using exons complementary to U5 Loop1, so all U5 pairs are Watson-Crick. In a real situation Watson-Crick pairs will be interspersed by isosteric pairs, that can achieve Watson-Crick-like geometry by tautomerisation (Artemyeva-Isman and Porter, 2021). The absolutely conserved non-Watson-Crick pair between the ends of the intron G+1 and G-1 was set in accordance with mutation analyses as tWW: bases interact by Watson-Crick edges, glycosidic bonds in trans conformation (2nd Westhof geometric family). The local orientation of RNA strands is parallel in this configuration (see text). Clockwise - A-C: Model projections generated in UCSF Chimera (see below Data availability statement for the link to the pdb file). D: Schematic of secondary structure. Human 5’ exon consensus shows clearer complementarity to U5 Loop1 for introns lacking the conserved +5G: U5 forms more Watson-Crick pairs with the 5’ exon to support U6 interaction with the start of the intron, which is weakened by the absence of the U6 C42=G+5 pair.
Figure 7. Structural model of the U5 Loop1 base-pairing with the exons and the non-Watson-Crick pair between the ends of the intron. The model was generated by SimRNA v2.0 (Moafinejad et al., 2024) using exons complementary to U5 Loop1, so all U5 pairs are Watson-Crick. In a real situation Watson-Crick pairs will be interspersed by isosteric pairs, that can achieve Watson-Crick-like geometry by tautomerisation (Artemyeva-Isman and Porter, 2021). The absolutely conserved non-Watson-Crick pair between the ends of the intron G+1 and G-1 was set in accordance with mutation analyses as tWW: bases interact by Watson-Crick edges, glycosidic bonds in trans conformation (2nd Westhof geometric family). The local orientation of RNA strands is parallel in this configuration (see text). Clockwise - A-C: Model projections generated in UCSF Chimera (see below Data availability statement for the link to the pdb file). D: Schematic of secondary structure. Human 5’ exon consensus shows clearer complementarity to U5 Loop1 for introns lacking the conserved +5G: U5 forms more Watson-Crick pairs with the 5’ exon to support U6 interaction with the start of the intron, which is weakened by the absence of the U6 C42=G+5 pair.
Preprints 233649 g003
1
Westhof geometric classification of RNA base-pairs describes 12 families of bp defined by the interacting edges of the bases (Watson-Crick W, Hoogsteen H or Sugar S) and the orientation of the glycosidic bonds relative to the H-bonds (cis c or trans t). Canonical Watson-Crick pairs appear in the 1st geometric family defined as cWW (Leontis et al., 2002; NAKB Nucleic Acid Knowledgebase, 2026 - online RNA bp catalogue)
2
At low SAM concentrations METTL16 stably binds MAT2A 3’ UTR and acts as a splicing factor, boosting the excision of the last intron. If SAM is high MATTL16 methylates and disassociates from MAT2A 3’UTR and the last intron is retained. (In bacteria a SAM-sensitive riboswitch controls SAM synthetase.)
3
100KGP - 100,000 Genomes Project, UK; PFMG2025 - Plan France Médecine Génomique 2025
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.