Genomics
CRISPR Constructs
The Scoville Splice lineup spans ten tiers, from 3 million to 13 million SHU. Each tier is produced by a specific combination of CRISPR constructs β gene knockouts, overexpression cassettes, and ESM2-designed point mutations β assembled into a tiered architecture where each additional construct unlocks the next SHU ceiling. In total, the platform comprises 25 constructs, 24 guide RNAs, and 31,911 base pairs of computationally designed DNA.
How CRISPR-Cas9 Works
CRISPR-Cas9 is a genome editing system adapted from bacterial immune defense. It consists of two components: the Cas9 protein, which acts as molecular scissors, and a guide RNA (gRNA), which directs Cas9 to a specific location in the genome. The guide RNA is a short sequence β typically 20 nucleotides β complementary to the target DNA. When the guide RNA binds its target, Cas9 creates a double-strand break at that precise location.
The cellβs DNA repair machinery then fixes the break. If no template is provided, the repair is error-prone (non-homologous end joining, or NHEJ), typically introducing small insertions or deletions that disrupt the gene β a knockout. If a repair template is provided (homology-directed repair, or HDR), the cell incorporates the template sequence, enabling precise edits such as point mutations or gene insertions. Scoville Splice uses both strategies: NHEJ for the POX knockout and HDR for overexpression cassettes and Pun1 point mutations.
Guide RNA Design
Each of the 24 guide RNAs was computationally designed to maximize on-target efficiency while minimizing off-target activity. The design process considers several factors:
- PAM site selection. Cas9 requires a protospacer adjacent motif (PAM) β the sequence NGG β immediately downstream of the target site. Guide RNAs are designed only at positions where a suitable PAM exists.
- On-target scoring. Computational models predict how efficiently a given guide RNA will direct Cas9 cutting. Guides with high predicted activity scores are prioritized.
- Off-target minimization. Each candidate guide is aligned against the entire Capsicum genome to identify potential off-target binding sites. Guides with close matches elsewhere in the genome are rejected or deprioritized.
- GC content. Guide RNAs with extreme GC content (very high or very low) tend to have reduced activity. The 24 selected guides fall within the optimal 40-70% GC range.
The Tiered Construct Architecture
The construct architecture is organized into tiers, with each successive tier adding one additional construct to the previous tierβs configuration. This modular approach is guided by metabolic flux analysis: the construct with the highest impact on the current bottleneck is added first, and each subsequent construct addresses the next limiting factor.
Tiers 1β4 β 3M to 5.1M SHU
1 construct: POX knockout
Peroxidase (POX) degrades capsaicinoids during fruit ripening. Knocking it out eliminates this degradation pathway, allowing capsaicinoids to accumulate to higher levels. The four lowest tiers differ in cultivar background and growth conditions, not in genetic architecture.
Tier 5 β 6M SHU
2 constructs: POX KO + PAL overexpression
PAL (phenylalanine ammonia-lyase) is the entry enzyme for the phenylpropanoid branch. Overexpressing PAL floods the vanillylamine supply chain with precursor, directly addressing the 90% flux control bottleneck identified through metabolic modeling.
Tier 6 β 7M SHU
3 constructs: POX KO + PAL OE + Pun1 S39L
With PAL overexpression increasing substrate flow, the Pun1 enzyme begins approaching its native catalytic capacity. The ESM2-designed S39L mutation (Log-Likelihood Ratio 2.828) improves substrate affinity at the binding pocket, ensuring Pun1 keeps pace with the amplified vanillylamine supply.
Tiers 7β10 β 8.2M to 13M SHU
4 constructs: POX KO + PAL OE + Pun1 mutations + COMT OE
COMT (caffeic acid O-methyltransferase) overexpression eliminates a secondary bottleneck that emerges mid-pathway when PAL overexpression pushes more flux through the vanillylamine branch than COMT can handle at native expression levels. All three Pun1 mutations (S39L, L345G, C175S) are active. The four highest tiers differ in expression tuning, promoter strength, and cultivar selection.
Computational Design First
Every construct in the Scoville Splice architecture was computationally designed, modeled, and validated in silico before any wet-lab synthesis. Guide RNAs were screened against the full Capsicum genome for off-target activity. Overexpression cassettes were codon-optimized for the target cultivar. Pun1 mutations were scored using ESM2βs protein language model. Metabolic flux simulations predicted the SHU outcome of each construct combination before a single plant was transformed.
This computational-first approach dramatically reduces the number of physical experiments required. Rather than screening thousands of random mutations, the design space is narrowed to high-confidence candidates before any DNA is synthesized. The 25 constructs represent the computationally optimal solution, not a sample from a larger library.
Public Record
The complete construct sequences have been deposited under GenBank accession SUB16548149. This public submission provides the research community with full access to the designed sequences and the tiered construct architecture used across the Scoville Splice platform.
By the Numbers
25
Constructs
24
Guide RNAs
31,911
Base Pairs