minibwa v0.5+0.6: * fixed an issue in PE mapQ (inherited from bwa-mem) * more conservative PE mapQ (learned from dragmap) * HG002 var calling: FP down but FN up * memory-mapped index loading * user-defined insert size distribution * APIs working for BS-seq github.com/lh3/minibwa/...
Профиль
Heng Li
Профиль VivelyAssociate Professor DFCI & HMS
minibwa-0.4 released with minor improvement and a few fixes to typos. Also a new blog post on "minibwa is the new bwa-mem": lh3.github.io/2026/07/04/m...
Minibwa is the new bwa-memlh3.github.iominibwa v0.3 released with a few minor bug fixes and two missing bwa-mem features (XA tag for secondary hits and option -H to inject header lines). Also added the "mem" subcommand to mimic "bwa mem" CLI to some extent. github.com/lh3/minibwa/...
minibwa-0.2 released with a few minor bug fixes (please keep bug reports coming). No algorithm changes. Minibwa is also available via bioconda thanks to Thanh Lee. github.com/lh3/minibwa/...
Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357
Jeremy Wang developed rammap, a minimap2 rewrite in Rust. It achieves comparable or better performance than minimap2 and produces identical output to minimap2. During rewrite, Jeremy found two long-existing bugs in minimap2 which are fixed in v2.31. www.biorxiv.org/content/10.6...
www.biorxiv.orgBlog post on "The AI Rewrite Dilemma": lh3.github.io/2026/04/17/t...
The AI Rewrite Dilemmalh3.github.ioLongcallR for competitive SNP calling and haplotype phasing, and simplified allele-specific analysis with long RNA-seq reads. Found ~100 junctions affected by SNPs per sample with most junctions novel. Developed by Neng Huang. Published in @natmethods.nature.com. Read at rdcu.be/faKhL
Long reads carry multiple small vars and SVs and their phasing. LongcallD is the only caller that tightly integrates germline/mosaic small/structural vars/MEIs and their phasing in a single C program. One command line to get competitive small variant calls and better SVs. Led by Yan Gao.
bioRxiv GenomicsLongcallD: joint calling and phasing of small, structural and mosaic variants from long reads https://www.biorxiv.org/content/10.64898/2026.03.20.713111v1
Pacific Biosciences Sells Short-Read Sequencing Assets to Illumina for $48.1M www.pacb.com/press_releas...
I am looking for a postdoc to develop high-performance algorithms in computational genomics. Email or DM me if interested. For more information, see hlilab.github.io/vacancies. RTs appreciated!
HLi Lab - VacanciesOpeningshlilab.github.ioNow published in Algorithms for Molecular Biology: link.springer.com/article/10.1.... Key message: a tiny CNN model with 7k parameters can capture main splice signals across vertebrates+insect and halves the minimap2 & miniprot junction error rate. I always use this new feature now.
Heng LiPreprint on "Improving spliced alignment by modeling splice sites with deep learning". It describes minisplice for modeling splice signals. Minimap2 and miniprot now optionally use the predicted scores to improve spliced alignment. arxiv.org/abs/2506.12986
Now published in gigascience: academic.oup.com/gigascience/.... Key messages: SVs are highly enriched in low-complexity/tandem-repeat regions and are harder to call. They behave differently from transposon insertions. Always stratify if you study SVs.
Validate Useracademic.oup.comHeng LiDo you know ~60% of human SVs fall in ~1% of GRCh38? See our new preprint: arxiv.org/abs/2509.23057 and the companion blog post on how we started this project and longdust: lh3.github.io/2025/09/29/o.... Work with Alvin Qin
To those who have open data at AWS cloud: you can use the S3 Bucket Browser to list files in your buckets. You can either 1) put the bucket name at lh3.github.io/s3bb/, 2) or copy this index.html github.com/lh3/s3bb/blo... to the root of a bucket. Generated by gemini.
S3 Bucket Browserlh3.github.ioThanks to the AWS Open Data program, this dataset along some derived data is also openly accessible via @AWSCloud at openhgl.s3.us-east-1.amazonaws.com/index.html
S3 Bucket Browseropenhgl.s3.us-east-1.amazonaws.comHeng Li579 high-quality human genomes from @humanpangenome.bsky.social, Arab Pangenome and individual papers (CHM13, CN1, KSA001, I002C, YAO and KOREF1). Sequences available in the AGC format (3.7GB) and FM-index in the ropebwt3 format (20.3GB). For details, see github.com/lh3/human-asm
579 high-quality human genomes from @humanpangenome.bsky.social, Arab Pangenome and individual papers (CHM13, CN1, KSA001, I002C, YAO and KOREF1). Sequences available in the AGC format (3.7GB) and FM-index in the ropebwt3 format (20.3GB). For details, see github.com/lh3/human-asm