Language Models
brightmart/nlp_chinese_corpus
ε€§θ§ζ¨‘δΈζθͺηΆθ―θ¨ε€ηθ―ζ Large Scale Chinese Corpus for NLP
ποΈ Text Classificationβ Question Answering𧬠Embeddingsπ§ Natural Language Processing
OtherMITupdated Feb 6, 2026
[ICLR 2024] DNABERT-2: Efficient Foundation Model and Benchmark for Multi-Species Genome
$ git clone https://github.com/MAGICS-LAB/DNABERT_2.gitε€§θ§ζ¨‘δΈζθͺηΆθ―θ¨ε€ηθ―ζ Large Scale Chinese Corpus for NLP
δΈζθ―θ¨ηθ§£ζ΅θ―εΊε Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection and Instruction-Aware Models for Conversational AI
μΈμ΄λͺ¨λΈμ νμ΅νκΈ° μν κ³΅κ° νκ΅μ΄ instruction datasetλ€μ λͺ¨μλμμ΅λλ€.
Genomic Pretrained Network - GPN, GPN-MSA, PhyloGPN, GPN-Star
GENA-LM is a transformer masked language model trained on human DNA sequence.
Data from GitHub Β· snapshot Sep 24, 2026