Changelog
v1.0.14 (2026-09-25)
Honour the genetic code an annotation declares (
transl_table) in protein extraction, model translation, the ORF search and stop completion, and run miniprot once per declared code. Four of the 13 human mitochondrial genes in the published CHM13 annotation were truncated by translating table-2 CDS with the standard code. Table 1 output is unchanged.Place a reference gene at a second target locus when miniprot finds it there and nothing else reaches it, which is what a whole-genome duplication produces (+690 verified genes on human to zebrafish).
--no-rescue-second-locusopts out.Count every class of dropped reference feature and report it once at the end of the run, and in
run_manifest.json.Bound the isoform rescue's in-flight prefetch with
--rescue-max-inflight(default 8192): 24.7 % less peak memory for 4.4 % wall.Fall back to in-process work when a worker pool cannot fork, instead of aborting the run.
Do not abort a lift because a CDS names a
Parentthe annotation does not declare; the check becomes fatal under--strict-gff.Report alternate-locus and patch contigs in the reference.
Speed up Step 7 by roughly a tenth, and stop re-deriving reference-side work.
Normalize sparse coding references into explicit hierarchies, preserve CDS attributes, record generated IDs and check supplied FASTA aliases.
Preserve genes and exon/CDS attributes when converting GTF with gffread. Conversion files are private to each run, with hashes, command and tool provenance recorded in the run manifest.
Repair direct-GTF self-parent relations and preserve transcript IDs during inference. Recognize Ensembl
gene_biotypeand apply strict GFF3 grammar checks after conversion.Honour
transl_except: read declared selenocysteine and other recoded stops through when scoring, calling variants and chaining, and rewrite the attribute into the lifted model's coordinates (it kept the reference's). All 25 human selenoprotein genes were mis-scored or truncated before; on GRCh38 to CHM13 their 53 transcripts now average 0.998 identity. A distant lift can also place selenoprotein genes it used to miss (seven on human to zebrafish). Ensembl/GENCODE selenocysteine rows (Selenocysteine,stop_codon_redefined_as_selenocysteine) are read the same way: on GENCODE v49 to CHM13 the 71 selenoprotein transcripts go from 0.66-0.68 to 0.995-0.997.Miniprot-derived models use their reference transcript's declared genetic code. An ambiguous sparse reference model (e.g. NCBI GenBank yeast Ty genes) is lifted as written instead of aborting the run;
--strict-gffkeeps it fatal.Split a CDS spanning an intron into exonic segments with transcript-order phases. Reject ambiguous overlaps (the transcript, not its gene), and validate CDS-within-exon containment.
Rebind a trans-spliced transcript in precomputed Liftoff annotations to its unique containing gene fragment, including on the same sequence (drosophila
mod(mdg4)); preserve that fragment'spart. An ambiguous family is left as bound and counted instead of aborting the run.Skip a miniprot candidate whose CDS cannot be split at exon boundaries instead of letting the error drop the competing Liftoff gene.
Keep a gene when one of its lifted transcripts encodes no protein; it was dropped whole in every release since v1.0.9.
Stop revisiting a gene's tRNA/rRNA children as loci of their own, which recorded hundreds of false pipeline failures on mammalian RefSeq lifts.
Do not abort the isoform rescue on an empty batch when
LIFTON_RESCUE_ISOFORM_WORKERSis set (affected v1.0.12 and v1.0.13).Keep GTF conversions out of the system temp directory in evaluation mode and when no converter is installed.
Count what the miniprot-only rescue abandons: it is on by default, abandons candidates in ~40 places and recorded none, so a drop total of 0 meant "nothing was counted". Deliberate filters stay uncounted; failed lookups do not.
Index the transcript of a reference gene that also declares its own exons (RefSeq's organellar convention), so a miniprot hit on it can be resolved.
Report
genes_emitted_without_childrenonly where the reference gave the gene children; the raw tally is nowbare_gene_lines_emitted. It was 11,130 reported across four genomes against 2 real.Make
--threads 1and--threads Nagree again: the per-locus runtime asked for a transcript's exons recursively while the Step-7 proxy cached the direct children. On RefSeq's organellar convention that returned each exon twice, so a default single-threaded rice lift emitted 17 genes with duplicated exons and a doubled CDS.Stop emitting overlapping exons within one transcript (the open half of GH #26): miniprot's redundant
stop_codonis no longer ingested as a second exon, and a rebuilt exon is reconciled against the one it ran into. Reference annotations have none of these; LiftOn emitted 27 on CHM13, 96 on human to zebrafish, 47 on drosophila, 30 on rice and 13 on bee.Add the overlapping-exon and overlapping-CDS checks
gff3-validate's documentation has always claimed, within one sequence and strand; a declared ribosomal-slippage overlap is a warning. The CDS phase check reads a 5'-partial model's own first phase.Index the windowed aligner's reference only over the query's k-mers: same anchors, same windows, 1.3-1.9x faster anchor construction.
Strengthen release evidence with candidate/reference roles, artifact receipts, explicit expected cells and rejection of incomplete or stale results.
Require campaign and attempt history, isolate imported source modules, and keep failed reference protein extraction in unresolved coding accounting.
Check second-locus rescue against assembly-matched target RefSeq coding models and seeded shifted-locus controls, retaining every baseline feature.
v1.0.13
Packaging release candidate (2026-09-18), addressing GH #78.
Standard pip installs exclude mappy;
lifton[native]or prebuilt Conda mappy supplies the optional, explicitly enabled experimental binding.Missing-aligner messages give an installation command. Pip installs Python packages, while fresh standard lifts require minimap2 and miniprot on PATH.
Correct pip/macOS/source instructions and document complete Seqera environments.
Qualify built wheel/sdist in compiler-free containers and execute fresh native lifts before publishing. Genuine optional binding tests remain separately tested.
Updated Bioconda recipe preparation removes cigar and the obsolete setuptools cap and retains the DuckDB exclusions, prebuilt mappy and both aligners.
v1.0.12
Resource-safety and diagnostics release (2026-09-14), motivated by GH #71's two
native SIGSEGV failures on an approximately 20-Gb target. It also recovers
far more genes between distantly related species and runs multi-threaded by
default.
Accuracy between distant species:
A second miniprot-only rescue sub-pass reconsiders hits the rescue's length band rejected. That band compares genomic spans, which include introns and therefore follow genome size: lifting from human into fish, bird, and frog genomes it discarded 75–83 % of the missed genes miniprot found, although their alignments covered the whole reference protein. The sub-pass gates on protein coverage (≥ 0.8) with the same identity floor. Whole-genome primary-assembly gene recall rises from 0.315 to 0.593 (human → zebrafish), 0.364 to 0.615 (human → chicken) and 0.361 to 0.638 (human → xenopus), with no gene lost, no duplicate model, and 99.7 % of added models overlapping an annotated CDS of the released target annotation.
--no-coverage-rescue-gateopts out.Rescued genes now carry the other reference transcripts whose miniprot hits lie at their locus, each held to the same identity floor, without changing gene placement.
--no-rescue-isoformsopts out.
Speed:
--threads Nwith N > 1 now fans out Steps 7 and 8 without--locus-pipeline; output is byte-identical to the serial path.--no-locus-pipelinerestores serial processing.Liftoff's alignment parsing and GFF3 writing are faster, and with
--threads Nits lift loop runs one forked worker per reference chromosome (--no-parallel-liftopts out). Output is identical; the fresh-Liftoff aligner phase at-t 8halves on drosophila (261 s to 134 s).Miniprot-derived models now carry their stop codon. miniprot's CDS ends at the last aligned codon, while the reference convention -- and every other model LiftOn emits -- includes the stop, and the ORF search could not add it because such a model has no UTR to search. Only 39-59 % of rescued models on the distant whole genomes ended in a stop. The terminal CDS and its exon now grow by those three bases when the genome has them and the reference protein ends in a stop, after the ORF search and only when the model does not get worse.
--no-orf-stop-completionopts out.-dir/--intermediate-dirputs a run's intermediate files, statistics, score table and manifest where you choose, so concurrent runs sharing an output directory no longer share onelifton_output/(GH #14).
Changed and fixed:
-copiesextra gene copies keep their transcripts, exons and CDS. Liftoff suffixes every feature of an extra copy with_<extra_copy_number>, so a copy arrives asgene-X_1/rna-X_1. LiftOn resolved the gene id back to the reference correctly but looked the transcript id up verbatim, so the lookup failed and the transcript was skipped -- taking every exon and CDS with it and leaving a gene line with no children. Across the 17-genome benchmark set this affected about 4,400 genes (rice 539 of 815 copy genes, human to zebrafish 1,178 of 1,881), identically in v1.0.11, so it was present in every release. The reference transcript is now resolved the way the gene already was: the exact id first, and only on a miss the_<N>copy base, accepted only when that base really is a transcript of the reference gene the copy belongs to -- so reference ids that genuinely end in_<int>are left alone.A gene emitted with no child features is counted and reported at the end of the run and in
run_manifest.jsonasgenes_emitted_without_children.A flat annotation -- a prokaryotic GFF3 with top-level
CDSrows and nogene-- now lifts those rows instead of selecting nothing and failing several steps later inside Liftoff, and an empty selection stops the run at once, naming the feature types the annotation contains (GH #37).run_manifest.jsonrecords where the miniprot-only rescue spends its time, split across candidate placement, the coverage sub-pass, and the isoform pass's prefetch, scoring and attachment.Targets above 4,000,000,000 bases run Liftoff/minimap2 before miniprot by default, preventing the two index-memory peaks from overlapping. Smaller targets retain concurrent execution.
--serial-alignersand--parallel-alignersare mutually exclusive force overrides.Liftoff creates no more Python workers than alignment tasks and divides the configured thread budget across those tasks. One target with
--threads 40therefore creates one worker whose minimap2 command receives 40 threads.run_manifest.jsonrecords target statistics, schedule/reason, exact aligner commands, last completed stages, bounded stderr tails, return codes, and POSIX signal names.-11is reported asSIGSEGV; OOM is not asserted without OS or scheduler evidence.miniprot older than 0.14 is rejected for a target sequence at least 2^31 bases, covering its known large-sequence limitation.
Trans-spliced copies with repeated logical parent IDs now bind children to the matching parent on the same sequence.
Intermediate Liftoff and miniprot files are written to
lifton_output/liftoff/andlifton_output/miniprot/again; v1.0.10 and v1.0.11 wrote them tolifton_outputliftoff/andlifton_outputminiprot/.
See Large-genome resource failures for diagnosis and recovery guidance.
v1.0.11
Single-fix release (2026-08-01). v1.0.10 could not build a database from some
annotations that earlier versions handled, aborting the run outright. If you lift
plant genomes, or any annotation produced with -copies, upgrade.
Fixed:
``UNIQUE constraint failed: features.id`` on annotations containing copy features.
gffutilsdisambiguates a repeatedcds-Xby renaming it tocds-X_1— exactly the suffix Liftoff's-copiesmode gives extra gene copies. Where that generated name already belonged to a real feature the insert failed and LiftOn exited. LiftOn now renames only ids that are both repeated and whose<id>_<n>already exists, leaving legitimate discontinuous CDS — which share an id by design — untouched. Measured across the benchmark corpus: v1.0.8 ingests 34 of 34 inputs, v1.0.10 ingested 31, v1.0.11 ingests 34. Same error class as GH #47, #12 and #7.A defect in the database-build fallback is no longer reported as a bad input file.
NameError,AttributeError,TypeErrorandImportErrorraised inside the fallback now propagate instead of being reported as a failed build.
v1.0.10
Incremental release (2026-07-30). Consolidates everything that accumulated
behind the v1.0.10 tag: the originally staged v1.0.10 work, a batch of
GitHub-issue fixes, a full-repository correctness audit, and a CDS-attribute
parity fix. Neither v1.0.10 nor the intermediate v1.0.10.1 was ever published,
so v1.0.9 is the baseline for every comparison below. As in v1.0.9, every
output-affecting default ships with an opt-out flag. See the project
CHANGELOG.md for the full list.
Changed (output-affecting defaults — results differ vs v1.0.9):
A third best-of-outcome merge candidate. Alongside {chained merge + ORF-rescue, Liftoff + ORF-rescue}, LiftOn now scores miniprot's native CDS-only model and adopts it only when its ORF-rescued protein identity is strictly higher than the two-way winner — so per-transcript identity never decreases — and never for an antisense hit. Pass
--no-miniprot-candidateto restore the two-way merge.A divergence-adaptive miniprot-only rescue floor. The rescue pass used a fixed protein-identity floor of 0.50; that floor is now lowered toward 0.30 as the DNA lift's gene recall drops, recovering more genuinely missing genes at large evolutionary distance while staying inert on same- and close-species lifts. Pass
--no-adaptive-rescue-floorto restore the fixed floor.Richer, spec-valid CDS rows. Rebuilt CDS lines lost the reference's descriptive attributes (
Dbxref,product,protein_id,gene,locus_tag, ...) and carried noIDat all. They now inherit those attributes and share oneID=cds-<transcript>— the spec's discontinuous-CDS form. Coordinates and the encoded protein are untouched; output grows 12–43%.LIFTON_NO_CDS_ATTR_CARRY=1andLIFTON_NO_CONTAINMENT_NORMALIZE=1reproduce the previous bytes.Coding transcripts are harmonized to ``mRNA``. The DNA-lift path preserved the reference featuretype — often the generic
transcript— while the miniprot path emittedmRNA, so one output labelled the same kind of feature two ways.LIFTON_NO_MRNA_HARMONIZE=1preserves the reference type.
Fixed (reported issues):
no such column: <id>from the DNA lift on strict SQLite builds (GH #35), which is why the same input worked on one machine and failed on another.A single unresolvable alignment killing an entire chromosome (GH #39).
Mixed-strand output from an antisense merge (GH #33).
UTR indels reported as frameshifts on close pairs (GH #46).
CDS emitted with no
ID(GH #32, GH #8) and coding transcripts typed inconsistently (GH #28).Header-only Liftoff output from an invalid or LFS-pointer minimap2 index (GH #57), and a
--streamingest crash on DuckDB 1.5.3/1.5.4 (GH #56).minimap2 is now preflight-checked at startup and documented as a requirement (GH #43, GH #11); the run summary is also written to
stats/summary.txt(GH #50).
Fixed (robustness):
A skipped locus no longer withholds the whole annotation. One bad gene in a 60,000-gene genome used to produce exit 2 and only a
*.partial.gff3. Per-locus failures now publish, reportpartial_success, and are recorded inrun_manifest.json;--strict-completenessrestores the old behaviour.Write-funnel crashes on gene-like/organellar children and on inverted (
start > end) coordinates, which aborted a whole-genome write.Three-level hierarchies (
gene → primary_transcript → miRNA → exon) are lifted instead of silently dropped, and LiftOn no longer rejects its own valid Ensembl/GENCODE-shaped output.
Added:
Auditable run manifests and transactional output —
lifton_output/run_manifest.jsonrecords sanitized arguments, SHA-256 input fingerprints, tool versions, phase timings, counts, validation and failures; the GFF3 is staged and published atomically only on success.Always-on structural output validation before publication, recorded in the run manifest and not bypassable by
--allow-partial-output.Content-addressed annotation caches that rebuild when their manifest no longer matches the source, parser settings, backend, or LiftOn version.
Performance (byte-neutral):
GFF3 validation is bounded — one contiguous top-level block at a time instead of the whole file: 5.64 GB → 267 MB on a full dog→cat lift (21× less memory, 1.9× faster), with identical reports.
Step 7 is ~21% faster on Drosophila at
-t 8(redundant leaf queries removed, vectorized identity counters, a pruning ordered gate), with byte-identical output and flat peak memory.Features are cloned rather than deep-copied (33.06 µs → 5.00 µs per rich CDS), and per-stage in-flight bounds (
--step7-max-inflightand siblings) cap the memory a parallel run can hold.
v1.0.9
Incremental release (2026-06-21). Turns on several accuracy- and
completeness-improving defaults, hardens LiftOn against whole-genome-abort
crashes on full RefSeq / cross-species genomes, and adds byte-identical
performance fast-paths plus validation tools. Some new defaults change the
output annotation relative to v1.0.8 — each ships with an opt-out flag that
restores the previous behaviour. See the project CHANGELOG.md for the full
list.
Changed (output-affecting defaults — results differ vs v1.0.8):
Gene-like lift is now the default. LiftOn auto-detects every reference top-level parent type with a transcript/exon hierarchy and lifts them all (pseudogenes,
ncRNA_gene, structured mobile elements, ...), not justgene. This adds features. Pass--gene-onlyto restore the oldgene-only lift (--lift-gene-likeis a kept no-op alias).Best-of-outcome Liftoff/miniprot merge is now the default. Per transcript, LiftOn keeps whichever of {chained merge + ORF-rescue, Liftoff + ORF-rescue} yields the higher emitted protein identity, avoiding merges that could silently frameshift downstream CDS. Pass
--legacy-mergeto restore the pre-promotion unconditional merge (--optimizeis a kept no-op alias).Banded / windowed alignment is now the default for all gene sizes. The aligner uses anchor-windowed alignment above ~2500 aa / 8000 nt (giant genes are always memory-bounded, so titin-scale transcripts no longer OOM): identity-exact on same-species lifts, mean-neutral cross-species, much faster and lighter on memory. Pass
--full-dp-alignto restore the exact giant-only full-DP path (--fast-alignis a kept no-op alias).Miniprot-only rescue is now default-ON. When the DNA lift misses a reference coding gene entirely (its miniprot mRNA overlaps no lifted gene locus), LiftOn emits the miniprot-only model, tagged
lifton_rescue=miniprot_only. It runs as a separate pass after the main lift closes, gated by a protein-identity floor with a dedup guard, recovering genuinely-missing genes at large evolutionary distance. Pass--no-miniprot-rescue(orLIFTON_MINIPROT_RESCUE=0) to opt out (--miniprot-rescueis a kept no-op alias).
Fixed (robustness / crashes):
Gene-like child double-lift crash on full RefSeq genomes. A gene-like feature that is a child of a gene was being enumerated again as a top-level locus, producing a duplicate FASTA key and crashing the run. Full Arabidopsis now completes (~99.9% of coding transcripts recovered, vs ~28% before the crash) and full rice likewise (~77% → ~99.9%). Top-level-only annotations are byte-identical.
Inverted-coordinate write crash. A single malformed transcript with
start > end(e.g. on the dog→cat lift) used to abort an entire ~60k-transcript genome during the write phase; such a feature is now skipped and logged, and the rest of the genome completes.A malformed feature no longer aborts a whole genome. The transcript writer now catches the project's validation exception so one bad feature is dropped and logged instead of propagating out of the parent write phase.
Whole-genome-abort hardening around
consume()/__str__plus a recursion-limit guard and a full traceback dump in the vendored-Liftoff call path, so deep/odd inputs fail loudly per-feature instead of silently killing the run.
Added (performance — byte-identical fast-paths, same output, faster/lighter):
--stream— pipe miniprot output straight into an in-memory database, skipping theminiprot.gff3disk round-trip and SQLite re-ingest.--inmemory-liftoff— feed Liftoff's lifted features to the database in-process, skipping theliftoff.gff3disk write and re-ingest.--threads N --locus-pipeline— fan out per-locus work across a thread pool; output is emitted in submission order so--threads Nis byte-identical to--threads 1(now works on the default backend without--native).--native— enable experimental native compatibility hooks. The mappy Liftoff route also requiresLIFTON_NATIVE_LIFTOFF_ALIGN=1; miniprot and bounded locus workers keep their proven paths.Concurrent aligner step is now the default (miniprot and Liftoff overlap); pass
--serial-alignersto opt out (--parallel-alignersis a kept no-op alias).Fused parallel Step 7 — per-locus materialise and process phases are fused into one pool, lowering both wall-clock and peak memory.
Sequence-extraction query collapse — collapses per-feature database queries (~2.2–2.4× fewer round-trips, ~25–34% faster extraction).
miniprot ``-t`` now scales with ``--threads`` (was pinned to its built-in default of 4; the default
-t 1is byte-identical).
Added (validation):
--strict-gff— run the NCBI GFF3 input-side validator on the reference annotation and exit non-zero on any spec violation.--validate-output/--validate-verbose— re-validate the just-written output GFF3 and print a structured report.``gff3-validate`` console script — a standalone GFF3 validator installed alongside
lifton.
Packaging:
Python floor raised to
>=3.10(thenetworkx>=3.3dependency requires Python ≥3.10; 3.9 is EOL).mappyis an optional dependency for the explicitly enabled native Liftoff route; the runtime falls back gracefully when it is absent.Added
MANIFEST.in,pyproject.toml(PEP 517/518 build config), and PyPI trove classifiers / project URLs.The vendored
gffbaseships its pure-Python fallback parser (no pre-built.so), so installs work without a Rust toolchain.
v1.0.7
Bug Fixes:
Fixed gffutils UNIQUE constraint errors: Enhanced duplicate feature ID handling with automatic recovery strategies. The system now automatically handles duplicate IDs in polished liftoff output and RefSeq annotations with non-overlapping CDS features, preventing silent failures during database creation.
Fixed GTF file format processing: Added automatic GTF format detection and proper handling. GTF files are now correctly processed with automatic gene/transcript inference, and optional automatic conversion to GFF3 format using gffread or agat tools.
Fixed ID parsing for IDs ending with numbers: Improved get_ID_base() function to safely handle feature IDs that naturally end with underscore and number (e.g., FMUND_1). The function now only removes suffixes when confirmed to be copy numbers, preventing silent failures.
Fixed CDS ID preservation: CDS features now preserve their IDs in output files, complying with GFF3 specification. CDS features from the same mRNA can now share the same ID as required by the standard.
Improvements:
Enhanced biotype attribute support: Added support for generic biotype attribute as fallback when gene_biotype (RefSeq) or gene_type (GENCODE/ENSEMBL/CHESS) are not present. This ensures protein-coding features are correctly identified regardless of annotation source.
Automatic GTF to GFF3 conversion: Added optional automatic conversion of GTF files to GFF3 format for better compatibility. Conversion uses gffread (preferred) or agat tools if available, with graceful fallback to direct GTF processing.
Improved error handling: Enhanced error messages and recovery strategies for database creation failures, providing better user guidance and automatic problem resolution.
Better format detection: Improved file format detection logic that checks multiple lines and patterns to reliably distinguish between GTF and GFF3 formats, with GFF3 as the safe default.
v1.0.0
Initial release of LiftOn
Release via the documentation (https://khchao.com/LiftOn)
Released via the paper (bioRxiv coming soon!)