๐Ÿ† #3,020 overall#138 of 148 in Tokenization & Preprocessing

bevacqua /megamark

:heart_eyes_cat: Markdown with easy tokenization, a fast highlighter, and a lean HTML sanitizer

$ git clone https://github.com/bevacqua/megamark.git
GitHub social preview for bevacqua/megamark
Stars
108
108
Forks
7
7
Language
JavaScript
License
MIT
Created
Feb 19, 2015
11.6 years old
Last push
Feb 18, 2021

Categories

GitHub topics

More in Tokenization & Preprocessing

Tokenization & Preprocessing

theseer/tokenizer

A small library for converting tokenized PHP source code into XML (and potentially other formats)

PHPOtherupdated Feb 3, 2026
GitHub โ†—โ˜… 5.2Kโ‘‚ 25
Tokenization & Preprocessing

marcelroed/gigatoken

Language model tokenization at GB/s

๐Ÿง  Natural Language Processing
RustMITupdated Sep 2, 2026
GitHub โ†—โ˜… 4.1Kโ‘‚ 220
Tokenization & Preprocessing

andialbrecht/sqlparse

A non-validating SQL parser module for Python

PythonBSD-3-Clauseupdated Aug 13, 2026
GitHub โ†—โ˜… 4Kโ‘‚ 755
Tokenization & Preprocessing

Chevrotain/chevrotain๐Ÿ”ฅ active

Parser Building Toolkit for JavaScript

TypeScriptApache-2.0updated Sep 24, 2026
GitHub โ†—โ˜… 2.8Kโ‘‚ 222
Tokenization & Preprocessing

dqbd/tiktokenizer

Online playground for OpenAPI tokenizers

TypeScriptMITupdated Apr 24, 2025
GitHub โ†—โ˜… 1.7Kโ‘‚ 187
Tokenization & Preprocessing

natasha/natasha

Solves basic Russian NLP tasks, API for lower level Natasha projects

๐Ÿท๏ธ Named Entity Recognition๐Ÿง  Natural Language Processing
PythonMITupdated Apr 13, 2026
GitHub โ†—โ˜… 1.4Kโ‘‚ 121

Data from GitHub ยท snapshot Sep 24, 2026