GeneSweep
Sweep reference genes across hundreds of genomes at once: presence, identity, coverage and mutations. Everything runs in your browser, no files are uploaded anywhere.
Each FASTA record is treated as a separate reference sequence (minimum 30 bp). Use an annotated gene/CDS FASTA for a whole-genome gene set, not a single whole-chromosome record. Total reference sequence length is limited to 12 Mb and any one reference record to 100,000 bp for browser safety. The 10,000-record limit is an input limit, not a guarantee that every large reference-by-isolate workload will finish quickly. Start with smaller batches and monitor runtime.
One exact genome file name per line, with extension (for example
isolate_01.fasta). Results, CSV and FASTA will follow this order. Files not in the list go last. Names in the list with no uploaded file appear as Missing.Advanced scoring and seed settings
Match must be positive, mismatch and gap negative. Shorter seeds find more distant genes but run slower.
X-drop: how far the score may fall below its best value before an extension stops. Higher tolerates longer bad stretches but is slower.
Min ungapped score: the score a gap-free stretch needs to be kept for full alignment. Lower is more sensitive but gives more false leads.
X-drop: how far the score may fall below its best value before an extension stops. Higher tolerates longer bad stretches but is slower.
Min ungapped score: the score a gap-free stretch needs to be kept for full alignment. Lower is more sensitive but gives more false leads.
| Genome | Gene | Status | Identity % | Coverage % | Contig | Start | End | Strand | Copies | Bitscore | E-value | Variants | Truncation | Pieces |
|---|