Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
S's picture

S PRO

aeon37
2 1 20
gono984's profile picture Quazim0t0's profile picture Talle2-X's profile picture
·

AI & ML interests

None yet

Recent Activity

reacted to wop's post with 🔥 5 days ago
https://huggingface.co/bench-labs/GCTokenizer-v1 , a multilingual tokenizer which does not require a training corpus https://huggingface.co/bench-labs developed **GCTokenizer-v1**, which is a multi-lingual tokenizer Available in four sizes: 32K, 65K, 131K and 262K tokens "S, M, L, XL" It utilizes an encoding scheme which allows it to handle characters in any language around the world General (multi lingual) Consensus (from multiple model tokenizers consensus) Tokenizer We included an implementation script too, built like BPE- it can encode arbitrary text, most of the time, efficiently
View all activity

Organizations

ML intern explorers's profile picture

aeon37 's models 6

aeon37/Kimi-K3

Image-Text-to-Text • Updated 20 days ago

aeon37/DSV4-f

Text Generation • 158B • Updated Jun 13 • 5

aeon37/DSV4-p

Text Generation • 862B • Updated Jun 13 • 7

aeon37/Nous-moe-10b-a1b-8k-wsd-lr3e4-1t

10B • Updated Apr 17 • 8

aeon37/DeepSeek-V2-Lite

16B • Updated Feb 15 • 5

aeon37/qwen2vl-72b-bnb-4bit

73B • Updated Jan 8 • 13
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs