PRIMERCAT · METHODS NOTE

Computational methods, parameters, and scope

Abstract

This document describes how PrimerCat selects reference sequences, generates candidates, performs database screens, and ranks outputs. Thresholds reflect the current implementation; scores compare candidates and are not probabilities of experimental success.

Document v1.3Updated 2026-09-04Applies to PrimerCat Web
01

Inputs and reference sequences

Template choice determines every downstream coordinate, candidate, and screening statement, so PrimerCat includes the selected accession and rationale in the result.

Gene-name mode

PrimerCat currently supports human and mouse. PrimerCat retrieves RefSeq mRNAs from NCBI Nucleotide, preferring a MANE Select record when present; otherwise it selects the NM_ coding transcript with the longest CDS. The working window extends the CDS by 200 bp on each side and is capped at 3,000 bp.1,2

The rule applies only to records retrieved for that request (up to 10 IDs; the first five GenBank records are read), not an exhaustive comparison of every isoform. Mouse records generally lack MANE Select and therefore rely more often on the longest-CDS fallback.

Custom-sequence mode

Whitespace is removed, text is uppercased, and characters are validated. qPCR/PCR accepts A, C, G, T, and N and rejects templates with >10% N. CRISPR scans the supplied DNA directly and does not infer its species, genome assembly, or locus.

02

qPCR primer design

The qPCR pipeline combines template resolution, Primer3 generation, backend-dependent screening, and a five-component heuristic rank.3,4

Computational workflow

01

Candidate generation

Primer3 generates up to 30 pairs and pre-ranks them by ascending pair penalty.

02

Candidate reduction

Only the top max(2 × requested return count, 10) pairs are screened; the final rank therefore applies to this subset, not every possible pair.

03

Paired-amplicon screen

Production human and mouse modes query both a version-pinned local genome and its matched RefSeq RNA collection. Species-filtered NCBI RefSeq RNA BLAST is used only when the local genome is unavailable.

04

Component status and rank

Results separate sequence parameters, specificity evidence, and exon-spanning status. The detail-only rank retains the five-component heuristic recorded with the frozen benchmark; failed screening contributes zero specificity points rather than treating unknown evidence as a pass.

Default design constraints

ParameterCurrent valueMethodological role
Primer length18–25 nt; optimum 20 ntPrimer3 candidate generation
Tm58–62 °C; optimum 60 °CPer-primer target range
GC40–60%Per-primer target range
Amplicon80–200 bpDefault RT-qPCR interval
HomopolymerMaximum 4 ntPRIMER_MAX_POLY_X
Structure ceilingsself-any 45 °C; self-end 35 °C; hairpin 24 °CPrimer3 thermodynamic ceilings; pair complementarity uses 45/35 °C

The three screening scopes are not equivalent

BackendSearch scopeCurrent evaluationDoes not establish
Local genome Bowtie2GRCh38.p14 or GRCm39End-to-end; retains hits with ≥85% coverage and ≤2 mismatches; up to 64 decision hits per end, then composes 50–5,000 bp products by orientation, distance, and target locus.Can reveal non-exonic hits, but remains limited by assembly version, thresholds, and the hit cap.
Local RefSeq RNA Bowtie2Accessioned RNA matched to the assembly annotation releaseUp to 128 decision hits per end; reports products on the selected transcript, same-gene isoforms, other genes, and unclassified transcripts separately.Covers only transcripts in that fixed annotation; it is not the sample's expression profile and omits unannotated transcripts.
NCBI RefSeq RNA BLAST fallbackSpecies-filtered reference transcriptsUsed only when the local genome is unavailable; short-query BLAST with up to five hits per primer and HSP coverage ≥85%.A return-set-limited transcript screen that cannot establish whole-genome uniqueness.

Production references are fixed to human GRCh38.p14 (GCF_000001405.40; RefSeq 2025-08 annotation) and mouse GRCm39 (GCF_000001635.27; RefSeq 2024-02 annotation).13,14,15,16 View the current production evidence snapshot →

03

Endpoint PCR primer design

Endpoint PCR uses a single DNA template supplied by the researcher. Presets only initialise the controls; the values the researcher submits define the run.3

Design

Primer3 uses 18–25 nt, GC 40–60%, no more than four identical consecutive bases, and the same thermodynamic ceilings as qPCR. Standard, colony-PCR, and high-fidelity presets start at 150–800, 200–1,500, and 500–3,000 bp.

Target interval

When a researcher supplies a 1-based closed interval, PrimerCat requires the returned amplicon to contain it and reports primer and amplicon template coordinates.

Annealing guidance

PrimerCat displays a simple starting estimate of lower primer Tm − 3 °C and a gradient of lower Tm − 5 to −1 °C. It is not fitted to a particular polymerase or buffer.

Optional specificity screen

PrimerCat searches each primer against species-filtered RefSeq genomic records in NCBI nt, then pairs compatible opposite-strand hits on the same accession.

Pair-screen criteria

  • Query coverage ≥85%
  • Total mismatches ≤3
  • Exact match across the final three 3′ nucleotides
  • Default amplicon window 50–5,000 bp
  • Up to 100 hits per primer; up to 10 paired records displayed

“One paired record” describes only the returned BLAST set. A truncation flag is shown when the hit cap is reached. This is not an exhaustive whole-genome in-silico PCR and common variants are not checked.4,6

04

CRISPR gRNA design

PrimerCat scans PAM candidates on both strands, computes sequence-feature activity and off-target screening separately, and then orders candidates by the off-target label followed by the activity heuristic.7,8

NucleasePAMCandidateImplementation
SpCas9NGG (3′)20 nt + PAMBoth-strand scan
SpCas9-NGNG (3′)20 nt + PAMRelaxed PAM does not imply wild-type-equivalent activity
Cas12aTTTV (5′)PAM + 20 ntV = A/C/G

Activity score

The SpCas9/SpCas9-NG score is inspired by Doench Rule Set 2 but the implementation uses a reduced set of positional weights plus GC bands, poly-T/poly-G, seed homopolymers, and terminal preferences. Cas12a uses a separate simplified rule set. This is not a complete reproduction of the published model and excludes chromatin accessibility, cell type, delivery, and expression system.7

PAM basis

NG recognition by SpCas9-NG and the T-rich Cas12a PAM are grounded in the corresponding nuclease studies; measured activity can still vary markedly across variants, cells, and targets.9,10

Off-target screen

BackendScopeCurrent implementation
Local Bowtie2Reference genome20-nt end-to-end alignment, ≤3 mismatches by default, up to 32 alignments; candidate loci are then checked for the appropriate canonical PAM. Optional target coordinates anchor the on-target locus.
NCBI nt BLAST fallbackSpecies-filtered ntShort query, up to 15 hits; gapped, short, and >4-mismatch HSPs are excluded. Without a usable locus, the strongest hit is heuristically treated as on-target.

Bowtie2/BLAST sequence-similarity screening does not predict cleavage and omits bulges, structural variants, sample-specific variants, and cell state. A low-risk label is a prioritisation aid, not a biosafety conclusion.8

05

BLAST sequence search

PrimerCat submits BLAST requests to NCBI and supports compatible combinations of blastn, blastp, blastx, and tblastn with nt, nr, RefSeq RNA, RefSeq protein, and Swiss-Prot.6

PrimerCat submits the researcher's selected E-value and hit count (1–50). For each database hit, PrimerCat displays only the highest-bit-score HSP with E-value, raw/bit score, identity, gaps, coordinates, and accession; displayed query/subject strings are capped at 300 characters. Researchers cannot use a local-alignment result alone to establish homology, function, phylogeny, or multiple-testing conclusions.

06

Method boundaries for laboratory references

We present solutions, protocols, and safety records as curated reference material; their evidence character differs from live computation or database search.

ModuleWhat the page providesUse boundary
Solution preparationDeterministic molarity, dilution, or percentage calculations; curated recipes scale linearly to final volume.pH, temperature, purity, hydrate state, and addition order remain governed by the primary source and laboratory SOP.
ProtocolsCommon workflows are structured by applicability, materials, steps, quality controls, and safety notes.Research reference only; not a standard operating procedure experimentally validated by PrimerCat.
Chemical safetyStatic records summarise representative PubChem LCSS/GHS information and link to sources.Not a live regulatory database. Form, concentration, mixture, and supplier alter classification; the container label and current SDS govern actual work.

PubChem aggregates chemical information from many contributors. PrimerCat uses representative records rather than presenting aggregated data as one permanent regulatory conclusion.12

REF

References

These sources document algorithms, databases, and validation frameworks. Citation does not imply that PrimerCat reproduces every part of a published model.

  1. [1]

    Morales J, et al. (2022). “A joint NCBI and EMBL-EBI transcript set for clinical genomics and research” Nature. 604:310–315. doi:10.1038/s41586-022-04558-8 ↗

  2. [2]

    O’Leary NA, et al. (2016). “Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation” Nucleic Acids Research. 44:D733–D745. doi:10.1093/nar/gkv1189 ↗

  3. [3]

    Untergasser A, et al. (2012). “Primer3—new capabilities and interfaces” Nucleic Acids Research. 40:e115. doi:10.1093/nar/gks596 ↗

  4. [4]

    Ye J, et al. (2012). “Primer-BLAST: a tool to design target-specific primers for polymerase chain reaction” BMC Bioinformatics. 13:134. doi:10.1186/1471-2105-13-134 ↗

  5. [5]

    Langmead B, Salzberg SL. (2012). “Fast gapped-read alignment with Bowtie 2” Nature Methods. 9:357–359. doi:10.1038/nmeth.1923 ↗

  6. [6]

    Camacho C, et al. (2009). “BLAST+: architecture and applications” BMC Bioinformatics. 10:421. doi:10.1186/1471-2105-10-421 ↗

  7. [7]

    Doench JG, et al. (2016). “Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9” Nature Biotechnology. 34:184–191. doi:10.1038/nbt.3437 ↗

  8. [8]

    Hsu PD, et al. (2013). “DNA targeting specificity of RNA-guided Cas9 nucleases” Nature Biotechnology. 31:827–832. doi:10.1038/nbt.2647 ↗

  9. [9]

    Nishimasu H, et al. (2018). “Engineered CRISPR-Cas9 nuclease with expanded targeting space” Science. 361:1259–1262. doi:10.1126/science.aas9129 ↗

  10. [10]

    Zetsche B, et al. (2015). “Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system” Cell. 163:759–771. doi:10.1016/j.cell.2015.09.038 ↗

  11. [11]

    Bustin SA, et al. (2025). “MIQE 2.0: Revision of the Minimum Information for Publication of Quantitative Real-Time PCR Experiments Guidelines” Clinical Chemistry. 71:634–651. doi:10.1093/clinchem/hvaf043 ↗

  12. [12]

    Kim S, et al. (2025). “PubChem 2025 update” Nucleic Acids Research. 53:D1516–D1525. doi:10.1093/nar/gkae1059 ↗

  13. [13]

    NCBI RefSeq (2025). “Homo sapiens genome assembly GRCh38.p14” NCBI Datasets. RefSeq assembly GCF_000001405.40. Source ↗

  14. [14]

    NCBI RefSeq (2025). “Homo sapiens Annotation Release GCF_000001405.40-RS_2025_08” NCBI Eukaryotic Genome Annotation. GRCh38.p14 RefSeq annotation report. Source ↗

  15. [15]

    NCBI RefSeq (2024). “Mus musculus genome assembly GRCm39” NCBI Datasets. RefSeq assembly GCF_000001635.27. Source ↗

  16. [16]

    NCBI RefSeq (2024). “Mus musculus Annotation Release GCF_000001635.27-RS_2024_02” NCBI Eukaryotic Genome Annotation. GRCm39 RefSeq annotation report. Source ↗

Reproducibility checklist

Save the method context with each result

To reproduce a design, researchers should record at least the input sequence or accession, species, parameters, screening backend, database scope, run date, and candidate sequences.