850 lines
89 KiB
JSON
850 lines
89 KiB
JSON
{
|
||
"version": "1",
|
||
"pip_version": "25.0.1",
|
||
"install": [
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/05/99/49ee85903dee060d9f08297b4a342e5e0bcfca2f027a07b4ee0a38ab13f9/faster_whisper-1.2.1-py3-none-any.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=79a66ad50688c0b794dd501dc340a736992a6342f7f95e5811be60b5224a26a7",
|
||
"hashes": {
|
||
"sha256": "79a66ad50688c0b794dd501dc340a736992a6342f7f95e5811be60b5224a26a7"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": true,
|
||
"metadata": {
|
||
"metadata_version": "2.1",
|
||
"name": "faster-whisper",
|
||
"version": "1.2.1",
|
||
"platform": [
|
||
"UNKNOWN"
|
||
],
|
||
"summary": "Faster Whisper transcription with CTranslate2",
|
||
"description": "[](https://github.com/SYSTRAN/faster-whisper/actions?query=workflow%3ACI) [](https://badge.fury.io/py/faster-whisper)\n\n# Faster Whisper transcription with CTranslate2\n\n**faster-whisper** is a reimplementation of OpenAI's Whisper model using [CTranslate2](https://github.com/OpenNMT/CTranslate2/), which is a fast inference engine for Transformer models.\n\nThis implementation is up to 4 times faster than [openai/whisper](https://github.com/openai/whisper) for the same accuracy while using less memory. The efficiency can be further improved with 8-bit quantization on both CPU and GPU.\n\n## Benchmark\n\n### Whisper\n\nFor reference, here's the time and memory usage that are required to transcribe [**13 minutes**](https://www.youtube.com/watch?v=0u7tTptBo9I) of audio using different implementations:\n\n* [openai/whisper](https://github.com/openai/whisper)@[v20240930](https://github.com/openai/whisper/tree/v20240930)\n* [whisper.cpp](https://github.com/ggerganov/whisper.cpp)@[v1.7.2](https://github.com/ggerganov/whisper.cpp/tree/v1.7.2)\n* [transformers](https://github.com/huggingface/transformers)@[v4.46.3](https://github.com/huggingface/transformers/tree/v4.46.3)\n* [faster-whisper](https://github.com/SYSTRAN/faster-whisper)@[v1.1.0](https://github.com/SYSTRAN/faster-whisper/tree/v1.1.0)\n\n### Large-v2 model on GPU\n\n| Implementation | Precision | Beam size | Time | VRAM Usage |\n| --- | --- | --- | --- | --- |\n| openai/whisper | fp16 | 5 | 2m23s | 4708MB |\n| whisper.cpp (Flash Attention) | fp16 | 5 | 1m05s | 4127MB |\n| transformers (SDPA)[^1] | fp16 | 5 | 1m52s | 4960MB |\n| faster-whisper | fp16 | 5 | 1m03s | 4525MB |\n| faster-whisper (`batch_size=8`) | fp16 | 5 | 17s | 6090MB |\n| faster-whisper | int8 | 5 | 59s | 2926MB |\n| faster-whisper (`batch_size=8`) | int8 | 5 | 16s | 4500MB |\n\n### distil-whisper-large-v3 model on GPU\n\n| Implementation | Precision | Beam size | Time | YT Commons WER |\n| --- | --- | --- | --- | --- |\n| transformers (SDPA) (`batch_size=16`) | fp16 | 5 | 46m12s | 14.801 |\n| faster-whisper (`batch_size=16`) | fp16 | 5 | 25m50s | 13.527 |\n\n*GPU Benchmarks are Executed with CUDA 12.4 on a NVIDIA RTX 3070 Ti 8GB.*\n[^1]: transformers OOM for any batch size > 1\n\n### Small model on CPU\n\n| Implementation | Precision | Beam size | Time | RAM Usage |\n| --- | --- | --- | --- | --- |\n| openai/whisper | fp32 | 5 | 6m58s | 2335MB |\n| whisper.cpp | fp32 | 5 | 2m05s | 1049MB |\n| whisper.cpp (OpenVINO) | fp32 | 5 | 1m45s | 1642MB |\n| faster-whisper | fp32 | 5 | 2m37s | 2257MB |\n| faster-whisper (`batch_size=8`) | fp32 | 5 | 1m06s | 4230MB |\n| faster-whisper | int8 | 5 | 1m42s | 1477MB |\n| faster-whisper (`batch_size=8`) | int8 | 5 | 51s | 3608MB |\n\n*Executed with 8 threads on an Intel Core i7-12700K.*\n\n\n## Requirements\n\n* Python 3.9 or greater\n\nUnlike openai-whisper, FFmpeg does **not** need to be installed on the system. The audio is decoded with the Python library [PyAV](https://github.com/PyAV-Org/PyAV) which bundles the FFmpeg libraries in its package.\n\n### GPU\n\nGPU execution requires the following NVIDIA libraries to be installed:\n\n* [cuBLAS for CUDA 12](https://developer.nvidia.com/cublas)\n* [cuDNN 9 for CUDA 12](https://developer.nvidia.com/cudnn)\n\n**Note**: The latest versions of `ctranslate2` only support CUDA 12 and cuDNN 9. For CUDA 11 and cuDNN 8, the current workaround is downgrading to the `3.24.0` version of `ctranslate2`, for CUDA 12 and cuDNN 8, downgrade to the `4.4.0` version of `ctranslate2`, (This can be done with `pip install --force-reinstall ctranslate2==4.4.0` or specifying the version in a `requirements.txt`).\n\nThere are multiple ways to install the NVIDIA libraries mentioned above. The recommended way is described in the official NVIDIA documentation, but we also suggest other installation methods below. \n\n<details>\n<summary>Other installation methods (click to expand)</summary>\n\n\n**Note:** For all these methods below, keep in mind the above note regarding CUDA versions. Depending on your setup, you may need to install the _CUDA 11_ versions of libraries that correspond to the CUDA 12 libraries listed in the instructions below.\n\n#### Use Docker\n\nThe libraries (cuBLAS, cuDNN) are installed in this official NVIDIA CUDA Docker images: `nvidia/cuda:12.3.2-cudnn9-runtime-ubuntu22.04`.\n\n#### Install with `pip` (Linux only)\n\nOn Linux these libraries can be installed with `pip`. Note that `LD_LIBRARY_PATH` must be set before launching Python.\n\n```bash\npip install nvidia-cublas-cu12 nvidia-cudnn-cu12==9.*\n\nexport LD_LIBRARY_PATH=`python3 -c 'import os; import nvidia.cublas.lib; import nvidia.cudnn.lib; print(os.path.dirname(nvidia.cublas.lib.__file__) + \":\" + os.path.dirname(nvidia.cudnn.lib.__file__))'`\n```\n\n#### Download the libraries from Purfview's repository (Windows & Linux)\n\nPurfview's [whisper-standalone-win](https://github.com/Purfview/whisper-standalone-win) provides the required NVIDIA libraries for Windows & Linux in a [single archive](https://github.com/Purfview/whisper-standalone-win/releases/tag/libs). Decompress the archive and place the libraries in a directory included in the `PATH`.\n\n</details>\n\n## Installation\n\nThe module can be installed from [PyPI](https://pypi.org/project/faster-whisper/):\n\n```bash\npip install faster-whisper\n```\n\n<details>\n<summary>Other installation methods (click to expand)</summary>\n\n### Install the master branch\n\n```bash\npip install --force-reinstall \"faster-whisper @ https://github.com/SYSTRAN/faster-whisper/archive/refs/heads/master.tar.gz\"\n```\n\n### Install a specific commit\n\n```bash\npip install --force-reinstall \"faster-whisper @ https://github.com/SYSTRAN/faster-whisper/archive/a4f1cc8f11433e454c3934442b5e1a4ed5e865c3.tar.gz\"\n```\n\n</details>\n\n## Usage\n\n### Faster-whisper\n\n```python\nfrom faster_whisper import WhisperModel\n\nmodel_size = \"large-v3\"\n\n# Run on GPU with FP16\nmodel = WhisperModel(model_size, device=\"cuda\", compute_type=\"float16\")\n\n# or run on GPU with INT8\n# model = WhisperModel(model_size, device=\"cuda\", compute_type=\"int8_float16\")\n# or run on CPU with INT8\n# model = WhisperModel(model_size, device=\"cpu\", compute_type=\"int8\")\n\nsegments, info = model.transcribe(\"audio.mp3\", beam_size=5)\n\nprint(\"Detected language '%s' with probability %f\" % (info.language, info.language_probability))\n\nfor segment in segments:\n print(\"[%.2fs -> %.2fs] %s\" % (segment.start, segment.end, segment.text))\n```\n\n**Warning:** `segments` is a *generator* so the transcription only starts when you iterate over it. The transcription can be run to completion by gathering the segments in a list or a `for` loop:\n\n```python\nsegments, _ = model.transcribe(\"audio.mp3\")\nsegments = list(segments) # The transcription will actually run here.\n```\n\n### Batched Transcription\nThe following code snippet illustrates how to run batched transcription on an example audio file. `BatchedInferencePipeline.transcribe` is a drop-in replacement for `WhisperModel.transcribe`\n\n```python\nfrom faster_whisper import WhisperModel, BatchedInferencePipeline\n\nmodel = WhisperModel(\"turbo\", device=\"cuda\", compute_type=\"float16\")\nbatched_model = BatchedInferencePipeline(model=model)\nsegments, info = batched_model.transcribe(\"audio.mp3\", batch_size=16)\n\nfor segment in segments:\n print(\"[%.2fs -> %.2fs] %s\" % (segment.start, segment.end, segment.text))\n```\n\n### Faster Distil-Whisper\n\nThe Distil-Whisper checkpoints are compatible with the Faster-Whisper package. In particular, the latest [distil-large-v3](https://huggingface.co/distil-whisper/distil-large-v3)\ncheckpoint is intrinsically designed to work with the Faster-Whisper transcription algorithm. The following code snippet \ndemonstrates how to run inference with distil-large-v3 on a specified audio file:\n\n```python\nfrom faster_whisper import WhisperModel\n\nmodel_size = \"distil-large-v3\"\n\nmodel = WhisperModel(model_size, device=\"cuda\", compute_type=\"float16\")\nsegments, info = model.transcribe(\"audio.mp3\", beam_size=5, language=\"en\", condition_on_previous_text=False)\n\nfor segment in segments:\n print(\"[%.2fs -> %.2fs] %s\" % (segment.start, segment.end, segment.text))\n```\n\nFor more information about the distil-large-v3 model, refer to the original [model card](https://huggingface.co/distil-whisper/distil-large-v3).\n\n### Word-level timestamps\n\n```python\nsegments, _ = model.transcribe(\"audio.mp3\", word_timestamps=True)\n\nfor segment in segments:\n for word in segment.words:\n print(\"[%.2fs -> %.2fs] %s\" % (word.start, word.end, word.word))\n```\n\n### VAD filter\n\nThe library integrates the [Silero VAD](https://github.com/snakers4/silero-vad) model to filter out parts of the audio without speech:\n\n```python\nsegments, _ = model.transcribe(\"audio.mp3\", vad_filter=True)\n```\n\nThe default behavior is conservative and only removes silence longer than 2 seconds. See the available VAD parameters and default values in the [source code](https://github.com/SYSTRAN/faster-whisper/blob/master/faster_whisper/vad.py). They can be customized with the dictionary argument `vad_parameters`:\n\n```python\nsegments, _ = model.transcribe(\n \"audio.mp3\",\n vad_filter=True,\n vad_parameters=dict(min_silence_duration_ms=500),\n)\n```\nVad filter is enabled by default for batched transcription.\n\n### Logging\n\nThe library logging level can be configured like this:\n\n```python\nimport logging\n\nlogging.basicConfig()\nlogging.getLogger(\"faster_whisper\").setLevel(logging.DEBUG)\n```\n\n### Going further\n\nSee more model and transcription options in the [`WhisperModel`](https://github.com/SYSTRAN/faster-whisper/blob/master/faster_whisper/transcribe.py) class implementation.\n\n## Community integrations\n\nHere is a non exhaustive list of open-source projects using faster-whisper. Feel free to add your project to the list!\n\n\n* [speaches](https://github.com/speaches-ai/speaches) is an OpenAI compatible server using `faster-whisper`. It's easily deployable with Docker, works with OpenAI SDKs/CLI, supports streaming, and live transcription.\n* [WhisperX](https://github.com/m-bain/whisperX) is an award-winning Python library that offers speaker diarization and accurate word-level timestamps using wav2vec2 alignment\n* [whisper-ctranslate2](https://github.com/Softcatala/whisper-ctranslate2) is a command line client based on faster-whisper and compatible with the original client from openai/whisper.\n* [whisper-diarize](https://github.com/MahmoudAshraf97/whisper-diarization) is a speaker diarization tool that is based on faster-whisper and NVIDIA NeMo.\n* [whisper-standalone-win](https://github.com/Purfview/whisper-standalone-win) Standalone CLI executables of faster-whisper for Windows, Linux & macOS. \n* [asr-sd-pipeline](https://github.com/hedrergudene/asr-sd-pipeline) provides a scalable, modular, end to end multi-speaker speech to text solution implemented using AzureML pipelines.\n* [Open-Lyrics](https://github.com/zh-plus/Open-Lyrics) is a Python library that transcribes voice files using faster-whisper, and translates/polishes the resulting text into `.lrc` files in the desired language using OpenAI-GPT.\n* [wscribe](https://github.com/geekodour/wscribe) is a flexible transcript generation tool supporting faster-whisper, it can export word level transcript and the exported transcript then can be edited with [wscribe-editor](https://github.com/geekodour/wscribe-editor)\n* [aTrain](https://github.com/BANDAS-Center/aTrain) is a graphical user interface implementation of faster-whisper developed at the BANDAS-Center at the University of Graz for transcription and diarization in Windows ([Windows Store App](https://apps.microsoft.com/detail/atrain/9N15Q44SZNS2)) and Linux.\n* [Whisper-Streaming](https://github.com/ufal/whisper_streaming) implements real-time mode for offline Whisper-like speech-to-text models with faster-whisper as the most recommended back-end. It implements a streaming policy with self-adaptive latency based on the actual source complexity, and demonstrates the state of the art.\n* [WhisperLive](https://github.com/collabora/WhisperLive) is a nearly-live implementation of OpenAI's Whisper which uses faster-whisper as the backend to transcribe audio in real-time.\n* [Faster-Whisper-Transcriber](https://github.com/BBC-Esq/ctranslate2-faster-whisper-transcriber) is a simple but reliable voice transcriber that provides a user-friendly interface.\n* [Open-dubbing](https://github.com/softcatala/open-dubbing) is open dubbing is an AI dubbing system which uses machine learning models to automatically translate and synchronize audio dialogue into different languages.\n* [Whisper-FastAPI](https://github.com/heimoshuiyu/whisper-fastapi) whisper-fastapi is a very simple script that provides an API backend compatible with OpenAI, HomeAssistant, and Konele (Android voice typing) formats.\n\n## Model conversion\n\nWhen loading a model from its size such as `WhisperModel(\"large-v3\")`, the corresponding CTranslate2 model is automatically downloaded from the [Hugging Face Hub](https://huggingface.co/Systran).\n\nWe also provide a script to convert any Whisper models compatible with the Transformers library. They could be the original OpenAI models or user fine-tuned models.\n\nFor example the command below converts the [original \"large-v3\" Whisper model](https://huggingface.co/openai/whisper-large-v3) and saves the weights in FP16:\n\n```bash\npip install transformers[torch]>=4.23\n\nct2-transformers-converter --model openai/whisper-large-v3 --output_dir whisper-large-v3-ct2\n--copy_files tokenizer.json preprocessor_config.json --quantization float16\n```\n\n* The option `--model` accepts a model name on the Hub or a path to a model directory.\n* If the option `--copy_files tokenizer.json` is not used, the tokenizer configuration is automatically downloaded when the model is loaded later.\n\nModels can also be converted from the code. See the [conversion API](https://opennmt.net/CTranslate2/python/ctranslate2.converters.TransformersConverter.html).\n\n### Load a converted model\n\n1. Directly load the model from a local directory:\n```python\nmodel = faster_whisper.WhisperModel(\"whisper-large-v3-ct2\")\n```\n\n2. [Upload your model to the Hugging Face Hub](https://huggingface.co/docs/transformers/model_sharing#upload-with-the-web-interface) and load it from its name:\n```python\nmodel = faster_whisper.WhisperModel(\"username/whisper-large-v3-ct2\")\n```\n\n## Comparing performance against other implementations\n\nIf you are comparing the performance against other Whisper implementations, you should make sure to run the comparison with similar settings. In particular:\n\n* Verify that the same transcription options are used, especially the same beam size. For example in openai/whisper, `model.transcribe` uses a default beam size of 1 but here we use a default beam size of 5.\n* Transcription speed is closely affected by the number of words in the transcript, so ensure that other implementations have a similar WER (Word Error Rate) to this one.\n* When running on CPU, make sure to set the same number of threads. Many frameworks will read the environment variable `OMP_NUM_THREADS`, which can be set when running your script:\n\n```bash\nOMP_NUM_THREADS=4 python3 my_script.py\n```\n\n\n",
|
||
"description_content_type": "text/markdown",
|
||
"keywords": [
|
||
"openai",
|
||
"whisper",
|
||
"speech",
|
||
"ctranslate2",
|
||
"inference",
|
||
"quantization",
|
||
"transformer"
|
||
],
|
||
"home_page": "https://github.com/SYSTRAN/faster-whisper",
|
||
"author": "Guillaume Klein",
|
||
"license": "MIT",
|
||
"license_file": [
|
||
"LICENSE"
|
||
],
|
||
"classifier": [
|
||
"Development Status :: 4 - Beta",
|
||
"Intended Audience :: Developers",
|
||
"Intended Audience :: Science/Research",
|
||
"License :: OSI Approved :: MIT License",
|
||
"Programming Language :: Python :: 3",
|
||
"Programming Language :: Python :: 3 :: Only",
|
||
"Programming Language :: Python :: 3.9",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Topic :: Scientific/Engineering :: Artificial Intelligence"
|
||
],
|
||
"requires_dist": [
|
||
"ctranslate2<5,>=4.0",
|
||
"huggingface-hub>=0.21",
|
||
"tokenizers<1,>=0.13",
|
||
"onnxruntime<2,>=1.14",
|
||
"av>=11",
|
||
"tqdm",
|
||
"transformers[torch]>=4.23; extra == \"conversion\"",
|
||
"black==23.*; extra == \"dev\"",
|
||
"flake8==6.*; extra == \"dev\"",
|
||
"isort==5.*; extra == \"dev\"",
|
||
"pytest==7.*; extra == \"dev\""
|
||
],
|
||
"requires_python": ">=3.9",
|
||
"provides_extra": [
|
||
"conversion",
|
||
"dev"
|
||
]
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/61/ae/736820e2000ec3e8ef3c491e8db7e3f7d24283e413b9fa7d5250333e2c7c/silk_python-0.2.7-cp312-cp312-win_amd64.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=8cac59e432e342b5234c7056acab4447abd8789938d557fd5346bd46c05ac668",
|
||
"hashes": {
|
||
"sha256": "8cac59e432e342b5234c7056acab4447abd8789938d557fd5346bd46c05ac668"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": true,
|
||
"metadata": {
|
||
"metadata_version": "2.4",
|
||
"name": "silk-python",
|
||
"version": "0.2.7",
|
||
"dynamic": [
|
||
"author",
|
||
"author-email",
|
||
"classifier",
|
||
"description",
|
||
"description-content-type",
|
||
"home-page",
|
||
"keywords",
|
||
"license",
|
||
"requires-dist",
|
||
"requires-python",
|
||
"summary"
|
||
],
|
||
"summary": "silk encode and decode",
|
||
"description": "<h1 align=\"center\"><i>✨ pysilk ✨ </i></h1>\r\n\r\n<h3 align=\"center\">The python binding for <a href=\"https://github.com/kn007/silk-v3-decoder\">silk-v3-decoder</a> </h3>\r\n\r\n[](https://pypi.org/project/silk-python/)\r\n\r\n\r\n\r\n\r\n\r\n\r\n## 安装\r\n```bash\r\npip install silk-python\r\n```\r\n\r\n\r\n## 使用\r\n- encode\r\n```python\r\nimport pysilk\r\n\r\nwith open(\"verybiginput.pcm\", \"rb\") as pcm, open(\"output.silk\", \"wb\") as silk:\r\n pysilk.encode(pcm, silk, 24000, 24000)\r\n```\r\n\r\n- decode\r\n\r\n```python\r\nimport pysilk\r\n\r\nwith open(\"verybiginput.silk\", \"rb\") as silk, open(\"output.pcm\", \"wb\") as pcm:\r\n pysilk.decode(silk, pcm, 24000)\r\n```\r\n\r\n## 支持功能\r\n- 接受任何二进制的```file-like object```,比如```BytesIO```,可以流式解码大文件\r\n- 包装了silk的全部C接口的参数,当然他们都有合理的默认值\r\n- 基于```Cython```, [关键部位](https://github.com/synodriver/pysilk/blob/stream/pysilk/silk.pxd#L43-L65) 内联C函数,高性能\r\n\r\n\r\n## 公开函数\r\n```python\r\nfrom typing import BinaryIO\r\n\r\ndef encode(input: BinaryIO, output: BinaryIO, sample_rate: int, bit_rate: int, max_internal_sample_rate: int = 24000, packet_loss_percentage: int = 0, complexity: int = 2, use_inband_fec: bool = False, use_dtx: bool = False, tencent: bool = True) -> None: ...\r\ndef decode(input: BinaryIO, output: BinaryIO, sample_rate: int, frame_size: int = 0, frames_per_packet: int = 1, more_internal_decoder_frames: bool = False, in_band_fec_offset: int = 0, loss: bool = False) -> None: ...\r\n```\r\n\r\n## 公开异常\r\n```python\r\nclass SilkError(Exception):\r\n pass\r\n```\r\n\r\n### ✨v0.2.0✨\r\n合并了[CFFI](https://github.com/synodriver/pysilk-cffi) 的工作\r\n\r\n### 本机编译\r\n```\r\npython -m pip install setuptools wheel cython cffi\r\ngit clone https://github.com/synodriver/pysilk\r\ncd pysilk\r\ngit submodule update --init --recursive\r\npython setup.py bdist_wheel --use-cython --use-cffi\r\n```\r\n\r\n### 后端选择\r\n默认由py实现决定,在cpython上自动选择cython后端,在pypy上自动选择cffi后端,使用```SILK_USE_CFFI```环境变量可以强制选择cffi\r\n",
|
||
"description_content_type": "text/markdown",
|
||
"keywords": [
|
||
"silk",
|
||
"encode",
|
||
"decode",
|
||
"pcm",
|
||
"audio"
|
||
],
|
||
"home_page": "https://github.com/synodriver/pysilk",
|
||
"author": "synodriver",
|
||
"author_email": "diguohuangjiajinweijun@gmail.com",
|
||
"license": "BSD",
|
||
"classifier": [
|
||
"Development Status :: 4 - Beta",
|
||
"Operating System :: OS Independent",
|
||
"License :: OSI Approved :: BSD License",
|
||
"Topic :: Multimedia :: Sound/Audio",
|
||
"Programming Language :: C",
|
||
"Programming Language :: Cython",
|
||
"Programming Language :: Python",
|
||
"Programming Language :: Python :: 3.8",
|
||
"Programming Language :: Python :: 3.9",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: Implementation :: CPython",
|
||
"Programming Language :: Python :: Implementation :: PyPy"
|
||
],
|
||
"requires_dist": [
|
||
"cffi>=1.0.0"
|
||
],
|
||
"requires_python": ">=3.6"
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/33/b4/76ba21e46704f632004276b85289a1582e95f5eff760436d6149875a1881/av-18.1.0-cp311-abi3-win_amd64.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=ea1480b7a8d5405cb5f382b344731bf125fd2c1c6fae3964f6c48595628387ff",
|
||
"hashes": {
|
||
"sha256": "ea1480b7a8d5405cb5f382b344731bf125fd2c1c6fae3964f6c48595628387ff"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": false,
|
||
"metadata": {
|
||
"metadata_version": "2.4",
|
||
"name": "av",
|
||
"version": "18.1.0",
|
||
"dynamic": [
|
||
"license-file"
|
||
],
|
||
"summary": "Pythonic bindings for FFmpeg's libraries.",
|
||
"description": "PyAV\r\n====\r\n\r\nPyAV is a Pythonic binding for the [FFmpeg][ffmpeg] libraries. We aim to provide all the power and control of the underlying library, but manage the gritty details as much as possible.\r\n\r\n---\r\n\r\n[![GitHub Test Status][github-tests-badge]][github-tests] [![Documentation][docs-badge]][docs] [![Python Package Index][pypi-badge]][pypi] [![Conda Forge][conda-badge]][conda]\r\n\r\nPyAV is for direct and precise access to your media via containers, streams, packets, codecs, and frames. It exposes a few transformations of that data, and helps you get your data to/from other packages (e.g. NumPy and Pillow).\r\n\r\nThis power does come with some responsibility as working with media is horrendously complicated and PyAV can't abstract it away or make all the best decisions for you. If the `ffmpeg` command does the job without you bending over backwards, PyAV is likely going to be more of a hindrance than a help.\r\n\r\nBut where you can't work without it, PyAV is a critical tool.\r\n\r\n\r\nInstallation\r\n------------\r\n\r\nPyAV requires Python 3.11 or later. Binary wheels are provided on [PyPI][pypi] for Linux, macOS, and Windows with FFmpeg bundled. You can install these wheels by running:\r\n\r\n```bash\r\npip install av\r\n```\r\n\r\nAnother way of installing PyAV is via [conda-forge][conda-forge]:\r\n\r\n```bash\r\nconda install av -c conda-forge\r\n```\r\n\r\nSee the [Conda install][conda-install] docs to get started with Miniconda.\r\n\r\n\r\nAlternative installation methods\r\n--------------------------------\r\n\r\nDue to the complexity of the dependencies, PyAV is not always the easiest Python package to install from source. This release supports FFmpeg 8.x. To build the source distribution against an existing FFmpeg installation on Linux or macOS, run:\r\n\r\n> [!WARNING]\r\n> FFmpeg's development files and `pkg-config` must be available on your system.\r\n\r\n```bash\r\npip install av --no-binary av\r\n```\r\n\r\n\r\nInstalling from source\r\n----------------------\r\n\r\nOn Linux or macOS, build PyAV from a Git checkout with:\r\n\r\n```bash\r\ngit clone https://github.com/PyAV-Org/PyAV.git\r\ncd PyAV\r\nsource scripts/activate.sh\r\n\r\n# Build ffmpeg from source. You can skip this step if ffmpeg 8.x is already installed.\r\n./scripts/build-deps\r\n\r\n# Build PyAV\r\nmake\r\n\r\n# Testing\r\nmake test\r\n\r\n# Install globally\r\ndeactivate\r\npip install .\r\n```\r\n\r\nOn Windows, use a Conda environment and the FFmpeg development files maintained by the PyAV project:\r\n\r\n```powershell\r\ngit clone https://github.com/PyAV-Org/PyAV.git\r\ncd PyAV\r\nconda create --name pyav-dev --channel conda-forge python=3.11 cython setuptools numpy pillow pytest\r\nconda activate pyav-dev\r\n$ffmpegDir = Join-Path $env:CONDA_PREFIX \"Library\"\r\npython scripts\\fetch-vendor.py --config-file scripts\\ffmpeg-latest.json $ffmpegDir\r\npython setup.py build_ext --inplace --ffmpeg-dir=$ffmpegDir\r\npython -m pytest\r\n```\r\n\r\n---\r\n\r\nHave fun, [read the docs][docs], [come chat with us][discuss], and good luck!\r\n\r\n\r\n\r\n[conda-badge]: https://img.shields.io/conda/vn/conda-forge/av.svg?colorB=CCB39A\r\n[conda]: https://anaconda.org/conda-forge/av\r\n[docs-badge]: https://img.shields.io/badge/docs-on%20pyav.basswood.io-blue.svg\r\n[docs]: https://pyav.basswood.io\r\n[pypi-badge]: https://img.shields.io/pypi/v/av.svg?colorB=CCB39A\r\n[pypi]: https://pypi.org/project/av\r\n[discuss]: https://github.com/PyAV-Org/PyAV/discussions\r\n\r\n[github-tests-badge]: https://github.com/PyAV-Org/PyAV/workflows/tests/badge.svg\r\n[github-tests]: https://github.com/PyAV-Org/PyAV/actions?workflow=tests\r\n[github]: https://github.com/PyAV-Org/PyAV\r\n\r\n[ffmpeg]: https://ffmpeg.org/\r\n[conda-forge]: https://conda-forge.github.io/\r\n[conda-install]: https://docs.conda.io/projects/conda/en/latest/user-guide/install/index.html\r\n",
|
||
"description_content_type": "text/markdown",
|
||
"author_email": "WyattBlue <wyattblue@auto-editor.com>, Jeremy Lainé <jeremy.laine@m4x.org>",
|
||
"license_expression": "BSD-3-Clause",
|
||
"license_file": [
|
||
"LICENSE.txt",
|
||
"AUTHORS.py",
|
||
"AUTHORS.rst"
|
||
],
|
||
"classifier": [
|
||
"Development Status :: 5 - Production/Stable",
|
||
"Intended Audience :: Developers",
|
||
"Natural Language :: English",
|
||
"Operating System :: MacOS :: MacOS X",
|
||
"Operating System :: POSIX",
|
||
"Operating System :: Unix",
|
||
"Operating System :: Microsoft :: Windows",
|
||
"Programming Language :: Cython",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: 3.14",
|
||
"Topic :: Software Development :: Libraries :: Python Modules",
|
||
"Topic :: Multimedia :: Sound/Audio",
|
||
"Topic :: Multimedia :: Sound/Audio :: Conversion",
|
||
"Topic :: Multimedia :: Video",
|
||
"Topic :: Multimedia :: Video :: Conversion"
|
||
],
|
||
"requires_python": ">=3.11",
|
||
"project_url": [
|
||
"Bug Tracker, https://github.com/PyAV-Org/PyAV/issues",
|
||
"Source Code, https://github.com/PyAV-Org/PyAV",
|
||
"homepage, https://pyav.basswood.io"
|
||
]
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/4e/23/e3b5322ff7368fcbed181ea4c209149416e7940b5b04971d5ee4084afe1a/ctranslate2-4.8.2-cp312-cp312-win_amd64.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=d94421d565d0de61c032998f737a18942b0f2bef40c0424b1846ec6f67300105",
|
||
"hashes": {
|
||
"sha256": "d94421d565d0de61c032998f737a18942b0f2bef40c0424b1846ec6f67300105"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": false,
|
||
"metadata": {
|
||
"metadata_version": "2.4",
|
||
"name": "ctranslate2",
|
||
"version": "4.8.2",
|
||
"dynamic": [
|
||
"author",
|
||
"classifier",
|
||
"description",
|
||
"description-content-type",
|
||
"home-page",
|
||
"keywords",
|
||
"license",
|
||
"project-url",
|
||
"requires-dist",
|
||
"requires-python",
|
||
"summary"
|
||
],
|
||
"summary": "Fast inference engine for Transformer models",
|
||
"description": "[](https://github.com/OpenNMT/CTranslate2/actions?query=workflow%3ACI) [](https://badge.fury.io/py/ctranslate2) [](https://opennmt.net/CTranslate2/) [](https://gitter.im/OpenNMT/CTranslate2?utm_source=badge&utm_medium=badge&utm_campaign=pr-badge) [](https://forum.opennmt.net/)\r\n\r\n# CTranslate2\r\n\r\nCTranslate2 is a C++ and Python library for efficient inference with Transformer models.\r\n\r\nThe project implements a custom runtime that applies many performance optimization techniques such as weights quantization, layers fusion, batch reordering, etc., to [accelerate and reduce the memory usage](#benchmarks) of Transformer models on CPU and GPU.\r\n\r\nThe following model types are currently supported:\r\n\r\n* Encoder-decoder models: Transformer base/big, M2M-100, NLLB, BART, mBART, Pegasus, T5, Whisper, T5Gemma, T5Gemma2, MADLAD-400\r\n* Decoder-only models: GPT-2, GPT-J, GPT-NeoX, OPT, BLOOM, MPT, Llama, Mistral, Gemma, CodeGen, GPTBigCode, Falcon, Qwen2\r\n* Encoder-only models: BERT, DistilBERT, XLM-RoBERTa\r\n\r\nCompatible models should be first converted into an optimized model format. The library includes converters for multiple frameworks:\r\n\r\n* [OpenNMT-py](https://opennmt.net/CTranslate2/guides/opennmt_py.html)\r\n* [OpenNMT-tf](https://opennmt.net/CTranslate2/guides/opennmt_tf.html)\r\n* [Fairseq](https://opennmt.net/CTranslate2/guides/fairseq.html)\r\n* [Marian](https://opennmt.net/CTranslate2/guides/marian.html)\r\n* [OPUS-MT](https://opennmt.net/CTranslate2/guides/opus_mt.html)\r\n* [Transformers](https://opennmt.net/CTranslate2/guides/transformers.html)\r\n\r\nThe project is production-oriented and comes with [backward compatibility guarantees](https://opennmt.net/CTranslate2/versioning.html), but it also includes experimental features related to model compression and inference acceleration.\r\n\r\n## Key features\r\n\r\n* **Fast and efficient execution on CPU and GPU**<br/>The execution [is significantly faster and requires less resources](#benchmarks) than general-purpose deep learning frameworks on supported models and tasks thanks to many advanced optimizations: layer fusion, padding removal, batch reordering, in-place operations, caching mechanism, etc.\r\n* **Quantization and reduced precision**<br/>The model serialization and computation support weights with [reduced precision](https://opennmt.net/CTranslate2/quantization.html): 16-bit floating points (FP16), 16-bit brain floating points (BF16), 16-bit integers (INT16), 8-bit integers (INT8) and AWQ quantization (INT4).\r\n* **Multiple CPU architectures support**<br/>The project supports x86-64 and AArch64/ARM64 processors and integrates multiple backends that are optimized for these platforms: [Intel MKL](https://software.intel.com/content/www/us/en/develop/tools/oneapi/components/onemkl.html), [oneDNN](https://github.com/oneapi-src/oneDNN), [OpenBLAS](https://www.openblas.net/), [Ruy](https://github.com/google/ruy), and [Apple Accelerate](https://developer.apple.com/documentation/accelerate).\r\n* **Automatic CPU detection and code dispatch**<br/>One binary can include multiple backends (e.g. Intel MKL and oneDNN) and instruction set architectures (e.g. AVX, AVX2) that are automatically selected at runtime based on the CPU information.\r\n* **Parallel and asynchronous execution**<br/>Multiple batches can be processed in parallel and asynchronously using multiple GPUs or CPU cores.\r\n* **Dynamic memory usage**<br/>The memory usage changes dynamically depending on the request size while still meeting performance requirements thanks to caching allocators on both CPU and GPU.\r\n* **Lightweight on disk**<br/>Quantization can make the models 4 times smaller on disk with minimal accuracy loss.\r\n* **Simple integration**<br/>The project has few dependencies and exposes simple APIs in [Python](https://opennmt.net/CTranslate2/python/overview.html) and C++ to cover most integration needs.\r\n* **Configurable and interactive decoding**<br/>[Advanced decoding features](https://opennmt.net/CTranslate2/decoding.html) allow autocompleting a partial sequence and returning alternatives at a specific location in the sequence.\r\n* **Support tensor parallelism for distributed inference**<br/>Very large model can be split into multiple GPUs. Following this [documentation](docs/parallel.md#model-and-tensor-parallelism) to set up the required environment.\r\n\r\nSome of these features are difficult to achieve with standard deep learning frameworks and are the motivation for this project.\r\n\r\n## Installation and usage\r\n\r\nCTranslate2 can be installed with pip:\r\n\r\n```bash\r\npip install ctranslate2\r\n```\r\n\r\nThe Python module is used to convert models and can translate or generate text with few lines of code:\r\n\r\n```python\r\ntranslator = ctranslate2.Translator(translation_model_path)\r\ntranslator.translate_batch(tokens)\r\n\r\ngenerator = ctranslate2.Generator(generation_model_path)\r\ngenerator.generate_batch(start_tokens)\r\n```\r\n\r\nSee the [documentation](https://opennmt.net/CTranslate2) for more information and examples.\r\n\r\nIf you have an AMD ROCm GPU, we provide specific Python wheels on the [releases page](https://github.com/OpenNMT/CTranslate2/releases/).\r\n\r\n## Web Server\r\n\r\n[ctranslate2-web-server](https://github.com/jordimas/ctranslate2-web-server) is a web server built on top of CTranslate2 that exposes an OpenAI-compatible REST API, making it easy to integrate CTranslate2 models into applications that already support the OpenAI API.\r\n\r\n## Benchmarks\r\n\r\nWe translate the En->De test set *newstest2014* with multiple models:\r\n\r\n* [OpenNMT-tf WMT14](https://opennmt.net/Models-tf/#translation): a base Transformer trained with OpenNMT-tf on the WMT14 dataset (4.5M lines)\r\n* [OpenNMT-py WMT14](https://opennmt.net/Models-py/#translation): a base Transformer trained with OpenNMT-py on the WMT14 dataset (4.5M lines)\r\n* [OPUS-MT](https://github.com/Helsinki-NLP/OPUS-MT-train/tree/master/models/en-de#opus-2020-02-26zip): a base Transformer trained with Marian on all OPUS data available on 2020-02-26 (81.9M lines)\r\n\r\nThe benchmark reports the number of target tokens generated per second (higher is better). The results are aggregated over multiple runs. See the [benchmark scripts](tools/benchmark) for more details and reproduce these numbers.\r\n\r\n**Please note that the results presented below are only valid for the configuration used during this benchmark: absolute and relative performance may change with different settings.**\r\n\r\n#### CPU\r\n\r\n| | Tokens per second | Max. memory | BLEU |\r\n| --- | --- | --- | --- |\r\n| **OpenNMT-tf WMT14 model** | | | |\r\n| OpenNMT-tf 2.31.0 (with TensorFlow 2.11.0) | 209.2 | 2653MB | 26.93 |\r\n| **OpenNMT-py WMT14 model** | | | |\r\n| OpenNMT-py 3.0.4 (with PyTorch 1.13.1) | 275.8 | 2012MB | 26.77 |\r\n| - int8 | 323.3 | 1359MB | 26.72 |\r\n| CTranslate2 3.6.0 | 658.8 | 849MB | 26.77 |\r\n| - int16 | 733.0 | 672MB | 26.82 |\r\n| - int8 | 860.2 | 529MB | 26.78 |\r\n| - int8 + vmap | 1126.2 | 598MB | 26.64 |\r\n| **OPUS-MT model** | | | |\r\n| Transformers 4.26.1 (with PyTorch 1.13.1) | 147.3 | 2332MB | 27.90 |\r\n| Marian 1.11.0 | 344.5 | 7605MB | 27.93 |\r\n| - int16 | 330.2 | 5901MB | 27.65 |\r\n| - int8 | 355.8 | 4763MB | 27.27 |\r\n| CTranslate2 3.6.0 | 525.0 | 721MB | 27.92 |\r\n| - int16 | 596.1 | 660MB | 27.53 |\r\n| - int8 | 696.1 | 516MB | 27.65 |\r\n\r\nExecuted with 4 threads on a [*c5.2xlarge*](https://aws.amazon.com/ec2/instance-types/c5/) Amazon EC2 instance equipped with an Intel(R) Xeon(R) Platinum 8275CL CPU.\r\n\r\n#### GPU\r\n\r\n| | Tokens per second | Max. GPU memory | Max. CPU memory | BLEU |\r\n| --- | --- | --- | --- | --- |\r\n| **OpenNMT-tf WMT14 model** | | | | |\r\n| OpenNMT-tf 2.31.0 (with TensorFlow 2.11.0) | 1483.5 | 3031MB | 3122MB | 26.94 |\r\n| **OpenNMT-py WMT14 model** | | | | |\r\n| OpenNMT-py 3.0.4 (with PyTorch 1.13.1) | 1795.2 | 2973MB | 3099MB | 26.77 |\r\n| FasterTransformer 5.3 | 6979.0 | 2402MB | 1131MB | 26.77 |\r\n| - float16 | 8592.5 | 1360MB | 1135MB | 26.80 |\r\n| CTranslate2 3.6.0 | 6634.7 | 1261MB | 953MB | 26.77 |\r\n| - int8 | 8567.2 | 1005MB | 807MB | 26.85 |\r\n| - float16 | 10990.7 | 941MB | 807MB | 26.77 |\r\n| - int8 + float16 | 8725.4 | 813MB | 800MB | 26.83 |\r\n| **OPUS-MT model** | | | | |\r\n| Transformers 4.26.1 (with PyTorch 1.13.1) | 1022.9 | 4097MB | 2109MB | 27.90 |\r\n| Marian 1.11.0 | 3241.0 | 3381MB | 2156MB | 27.92 |\r\n| - float16 | 3962.4 | 3239MB | 1976MB | 27.94 |\r\n| CTranslate2 3.6.0 | 5876.4 | 1197MB | 754MB | 27.92 |\r\n| - int8 | 7521.9 | 1005MB | 792MB | 27.79 |\r\n| - float16 | 9296.7 | 909MB | 814MB | 27.90 |\r\n| - int8 + float16 | 8362.7 | 813MB | 766MB | 27.90 |\r\n\r\nExecuted with CUDA 11 on a [*g5.xlarge*](https://aws.amazon.com/ec2/instance-types/g5/) Amazon EC2 instance equipped with a NVIDIA A10G GPU (driver version: 510.47.03).\r\n\r\n## Contributing\r\n\r\nCTranslate2 is a community-driven project. We welcome contributions of all kinds:\r\n* **New Model Support:** Help us implement more Transformer architectures.\r\n* **Performance:** Propose optimizations for CPU or GPU kernels.\r\n* **Bug Reports:** Open an issue if you find something not working as expected.\r\n* **Documentation:** Improve our guides or add new examples.\r\n\r\nCheck out our [Contributing Guide](CONTRIBUTING.md) to learn how to set up your development environment.\r\n\r\n## Additional resources\r\n\r\n* [Documentation](https://opennmt.net/CTranslate2)\r\n* [Forum](https://forum.opennmt.net)\r\n* [Gitter](https://gitter.im/OpenNMT/CTranslate2)\r\n",
|
||
"description_content_type": "text/markdown",
|
||
"keywords": [
|
||
"opennmt",
|
||
"nmt",
|
||
"neural",
|
||
"machine",
|
||
"translation",
|
||
"cuda",
|
||
"mkl",
|
||
"inference",
|
||
"quantization"
|
||
],
|
||
"home_page": "https://opennmt.net",
|
||
"author": "OpenNMT",
|
||
"license": "MIT",
|
||
"classifier": [
|
||
"Development Status :: 5 - Production/Stable",
|
||
"Environment :: GPU :: NVIDIA CUDA :: 12 :: 12.4",
|
||
"Intended Audience :: Developers",
|
||
"Intended Audience :: Science/Research",
|
||
"Programming Language :: Python :: 3",
|
||
"Programming Language :: Python :: 3 :: Only",
|
||
"Programming Language :: Python :: 3.9",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: 3.14",
|
||
"Topic :: Scientific/Engineering :: Artificial Intelligence"
|
||
],
|
||
"requires_dist": [
|
||
"numpy",
|
||
"pyyaml<7,>=5.3"
|
||
],
|
||
"requires_python": ">=3.9",
|
||
"project_url": [
|
||
"Documentation, https://opennmt.net/CTranslate2",
|
||
"Forum, https://forum.opennmt.net",
|
||
"Gitter, https://gitter.im/OpenNMT/CTranslate2",
|
||
"Source, https://github.com/OpenNMT/CTranslate2"
|
||
]
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/1b/cf/d98dd561d6d0d7b7d7a64d1563f8aaaa7c235daee41c1c9bcc3da62420ed/huggingface_hub-1.32.0-py3-none-any.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=b0c7c80561969d9cdacdd55fce67ba9584cca0b9d4ea80957a3a5c1445fac5c8",
|
||
"hashes": {
|
||
"sha256": "b0c7c80561969d9cdacdd55fce67ba9584cca0b9d4ea80957a3a5c1445fac5c8"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": false,
|
||
"metadata": {
|
||
"metadata_version": "2.4",
|
||
"name": "huggingface_hub",
|
||
"version": "1.32.0",
|
||
"dynamic": [
|
||
"author",
|
||
"author-email",
|
||
"classifier",
|
||
"description",
|
||
"description-content-type",
|
||
"home-page",
|
||
"keywords",
|
||
"license",
|
||
"license-file",
|
||
"provides-extra",
|
||
"requires-dist",
|
||
"requires-python",
|
||
"summary"
|
||
],
|
||
"summary": "Client library to download and publish models, datasets and other repos on the huggingface.co hub",
|
||
"description": "<p align=\"center\">\n <picture>\n <source media=\"(prefers-color-scheme: dark)\" srcset=\"https://huggingface.co/datasets/huggingface/documentation-images/raw/main/huggingface_hub-dark.svg\">\n <source media=\"(prefers-color-scheme: light)\" srcset=\"https://huggingface.co/datasets/huggingface/documentation-images/raw/main/huggingface_hub.svg\">\n <img alt=\"huggingface_hub library logo\" src=\"https://huggingface.co/datasets/huggingface/documentation-images/raw/main/huggingface_hub.svg\" width=\"352\" height=\"59\" style=\"max-width: 100%\">\n </picture>\n <br/>\n <br/>\n</p> \n\n<p align=\"center\">\n <i>The official CLI and Python client for the Hugging Face Hub.</i>\n <br/>\n <a href=\"#what-is-huggingface_hub\">About</a>\n ·\n <a href=\"https://huggingface.co/docs/huggingface_hub\">Documentation</a>\n ·\n <a href=\"https://huggingface.co/docs/huggingface_hub/en/installation\">Install</a>\n ·\n <a href=\"https://huggingface.co/docs/huggingface_hub/en/guides/cli\">CLI Guide</a>\n ·\n <a href=\"https://github.com/huggingface/huggingface_hub/blob/main/CONTRIBUTING.md\">Contributing</a>\n</p>\n\n<p align=\"center\">\n <a href=\"https://huggingface.co/docs/huggingface_hub/en/index\"><img alt=\"Documentation\" src=\"https://img.shields.io/website/http/huggingface.co/docs/huggingface_hub/index.svg?down_color=red&down_message=offline&up_message=online&label=doc\"></a>\n <a href=\"https://github.com/huggingface/huggingface_hub/releases\"><img alt=\"GitHub release\" src=\"https://img.shields.io/github/release/huggingface/huggingface_hub.svg\"></a>\n <a href=\"https://github.com/huggingface/huggingface_hub\"><img alt=\"PyPi version\" src=\"https://img.shields.io/pypi/pyversions/huggingface_hub.svg\"></a>\n <a href=\"https://pypi.org/project/huggingface-hub\"><img alt=\"PyPI - Downloads\" src=\"https://img.shields.io/pypi/dm/huggingface_hub\"></a>\n <a href=\"https://codecov.io/gh/huggingface/huggingface_hub\"><img alt=\"Code coverage\" src=\"https://codecov.io/gh/huggingface/huggingface_hub/branch/main/graph/badge.svg?token=RXP95LE2XL\"></a>\n</p>\n\n<h4 align=\"center\">\n <p>\n <b>English</b> |\n <a href=\"https://github.com/huggingface/huggingface_hub/blob/main/i18n/README_de.md\">Deutsch</a> |\n <a href=\"https://github.com/huggingface/huggingface_hub/blob/main/i18n/README_fr.md\">Français</a> |\n <a href=\"https://github.com/huggingface/huggingface_hub/blob/main/i18n/README_hi.md\">हिंदी</a> |\n <a href=\"https://github.com/huggingface/huggingface_hub/blob/main/i18n/README_ko.md\">한국어</a> |\n <a href=\"https://github.com/huggingface/huggingface_hub/blob/main/i18n/README_cn.md\">中文 (简体)</a> |\n <a href=\"https://github.com/huggingface/huggingface_hub/blob/main/i18n/README_kn.md\">ಕನ್ನಡ</a>\n </p>\n</h4>\n\n## Quick start\n\nInstall the [`hf` CLI](https://huggingface.co/docs/huggingface_hub/en/guides/cli) with the standalone installer:\n\n```bash\n# On macOS and Linux.\ncurl -LsSf https://hf.co/cli/install.sh | bash\n```\n\n```powershell\n# On Windows.\npowershell -ExecutionPolicy ByPass -c \"irm https://hf.co/cli/install.ps1 | iex\"\n```\n\nLog in, then start working with the Hub:\n\n```bash\n# Log in (use --token $HF_TOKEN in non-interactive environments)\nhf auth login\n\n# Find models served by Inference Providers\nhf models ls --warm\n\n# Download a model\nhf download Qwen/Qwen3-0.6B\n\n# Upload files to your own repo\nhf upload username/my-cool-model ./model.safetensors\n\n# Sync a local folder to a storage bucket\nhf buckets sync ./checkpoints hf://buckets/username/my-bucket\n\n# Run a job on Hugging Face infrastructure\nhf jobs run python:3.12 python -c \"print('Hello from the cloud!')\"\n\n# Discover everything else\nhf --help\n```\n\nThe Hub uses tokens to authenticate applications (see [docs](https://huggingface.co/docs/hub/security-tokens)). Check out the [CLI guide](https://huggingface.co/docs/huggingface_hub/en/guides/cli) for a tour of the main features.\n\n## What is `huggingface_hub`?\n\nThe `huggingface_hub` library allows you to interact with the [Hugging Face Hub](https://huggingface.co/), a platform democratizing open-source Machine Learning for creators and collaborators. Discover pre-trained models and datasets for your projects, play with the thousands of machine learning apps hosted on the Hub, or create and share your own models, datasets and demos with the community. Everything ships in one package with two interfaces: the [`hf` CLI](https://huggingface.co/docs/huggingface_hub/en/guides/cli) for your terminal and the `huggingface_hub` library for Python — both designed to work well for humans and AI agents. Use them to:\n\n- [Download files](https://huggingface.co/docs/huggingface_hub/en/guides/download) from the Hub.\n- [Upload files](https://huggingface.co/docs/huggingface_hub/en/guides/upload) to the Hub.\n- [Manage your repositories](https://huggingface.co/docs/huggingface_hub/en/guides/repository).\n- [Run Inference](https://huggingface.co/docs/huggingface_hub/en/guides/inference) on deployed models.\n- [Run Jobs](https://huggingface.co/docs/huggingface_hub/en/guides/jobs) on Hugging Face infrastructure.\n- [Search](https://huggingface.co/docs/huggingface_hub/en/guides/search) for models, datasets and Spaces.\n- [Share Model Cards](https://huggingface.co/docs/huggingface_hub/en/guides/model-cards) to document your models.\n- [Engage with the community](https://huggingface.co/docs/huggingface_hub/en/guides/community) through PRs and comments.\n- Do all of the above from the terminal with the [`hf` CLI](https://huggingface.co/docs/huggingface_hub/en/guides/cli).\n\n## Built for humans and AI agents\n\nThe `hf` CLI is designed for people and coding agents alike: the same commands adapt their output when run by an agent. If you use Claude Code, Codex, Cursor, or another coding agent, install the `hf` CLI Skill — a command reference generated from your installed CLI:\n\n```bash\n# for Codex, Cursor, OpenCode, Pi and other agents that load skills from `.agents/skills`\nhf skills add\n# includes the above + Claude Code\nhf skills add --claude\n```\n\nLearn more in the [Hugging Face CLI for AI agents guide](https://huggingface.co/docs/hub/agents-cli) and the [announcement blog post](https://huggingface.co/blog/hf-cli-for-agents).\n\n## Use the Python library\n\nInstall the `huggingface_hub` package with [pip](https://pypi.org/project/huggingface-hub/) (this also installs the `hf` CLI):\n\n```bash\npip install huggingface_hub\n```\n\nWe recommend using [`uv`](https://docs.astral.sh/uv/) for a fast and reliable install:\n\n```bash\nuv pip install huggingface_hub\n```\n\nIn order to keep the package minimal by default, `huggingface_hub` comes with optional dependencies useful for some use cases. For example, if you want to use the MCP module, run:\n\n```bash\npip install \"huggingface_hub[mcp]\"\n```\n\nTo learn more about installation and optional dependencies, check out the [installation guide](https://huggingface.co/docs/huggingface_hub/en/installation).\n\n### Download files\n\nDownload a single file\n\n```py\nfrom huggingface_hub import hf_hub_download\n\nhf_hub_download(repo_id=\"zai-org/GLM-5.2\", filename=\"config.json\")\n```\n\nOr an entire repository\n\n```py\nfrom huggingface_hub import snapshot_download\n\nsnapshot_download(\"sentence-transformers/all-MiniLM-L6-v2\")\n```\n\nFiles will be downloaded in a local cache folder. More details in [this guide](https://huggingface.co/docs/huggingface_hub/en/guides/manage-cache).\n\n### Create a repository\n\n```py\nfrom huggingface_hub import create_repo\n\ncreate_repo(repo_id=\"super-cool-model\")\n```\n\n### Upload files\n\nUpload a single file\n\n```py\nfrom huggingface_hub import upload_file\n\nupload_file(\n path_or_fileobj=\"/home/lysandre/dummy-test/README.md\",\n path_in_repo=\"README.md\",\n repo_id=\"lysandre/test-model\",\n)\n```\n\nOr an entire folder\n\n```py\nfrom huggingface_hub import upload_folder\n\nupload_folder(\n folder_path=\"/path/to/local/space\",\n repo_id=\"username/my-cool-space\",\n repo_type=\"space\",\n)\n```\n\nMore details in the [upload guide](https://huggingface.co/docs/huggingface_hub/en/guides/upload).\n\n## Integrating with the Hub.\n\nWe're partnering with cool open source ML libraries to provide free model hosting and versioning. You can find the existing integrations [here](https://huggingface.co/docs/hub/libraries).\n\nThe advantages are:\n\n- Free model or dataset hosting for libraries and their users.\n- Built-in file versioning, even with very large files, made possible by [Xet](https://huggingface.co/docs/hub/xet/index), the Hub's chunk-deduplicated storage backend.\n- In-browser widgets to play with the uploaded models.\n- Anyone can upload a new model for your library, they just need to add the corresponding tag for the model to be discoverable.\n- Fast downloads! We use Cloudfront (a CDN) to geo-replicate downloads so they're blazing fast from anywhere on the globe.\n- Usage stats and more features to come.\n\nIf you would like to integrate your library, feel free to open an issue to begin the discussion. We wrote a [step-by-step guide](https://huggingface.co/docs/hub/adding-a-library) with ❤️ showing how to do this integration.\n\n## Contributions (feature requests, bugs, etc.) are super welcome 💙💚💛💜🧡❤️\n\nEveryone is welcome to contribute, and we value everybody's contribution. Code is not the only way to help the community.\nAnswering questions, helping others, reaching out and improving the documentations are immensely valuable to the community.\nWe wrote a [contribution guide](https://github.com/huggingface/huggingface_hub/blob/main/CONTRIBUTING.md) to summarize\nhow to get started to contribute to this repository.\n",
|
||
"description_content_type": "text/markdown",
|
||
"keywords": [
|
||
"model-hub",
|
||
"machine-learning",
|
||
"models",
|
||
"natural-language-processing",
|
||
"deep-learning",
|
||
"pytorch",
|
||
"pretrained-models"
|
||
],
|
||
"home_page": "https://github.com/huggingface/huggingface_hub",
|
||
"author": "Hugging Face, Inc.",
|
||
"author_email": "julien@huggingface.co",
|
||
"license": "Apache-2.0",
|
||
"license_file": [
|
||
"LICENSE"
|
||
],
|
||
"classifier": [
|
||
"Intended Audience :: Developers",
|
||
"Intended Audience :: Education",
|
||
"Intended Audience :: Science/Research",
|
||
"License :: OSI Approved :: Apache Software License",
|
||
"Operating System :: OS Independent",
|
||
"Programming Language :: Python :: 3",
|
||
"Programming Language :: Python :: 3 :: Only",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: 3.14",
|
||
"Topic :: Scientific/Engineering :: Artificial Intelligence"
|
||
],
|
||
"requires_dist": [
|
||
"click<9.0.0,>=8.4.2",
|
||
"filelock>=3.10.0",
|
||
"fsspec>=2023.5.0",
|
||
"hf-xet<2.0.0,>=1.5.2; platform_machine == \"x86_64\" or platform_machine == \"amd64\" or platform_machine == \"AMD64\" or platform_machine == \"arm64\" or platform_machine == \"aarch64\"",
|
||
"httpx<1,>=0.23.0",
|
||
"packaging>=20.9",
|
||
"pyyaml>=5.1",
|
||
"tomli>=1.1.0; python_version < \"3.11\"",
|
||
"tqdm>=4.42.1",
|
||
"typing-extensions>=4.1.0",
|
||
"authlib>=1.3.2; extra == \"oauth\"",
|
||
"fastapi; extra == \"oauth\"",
|
||
"httpx; extra == \"oauth\"",
|
||
"itsdangerous; extra == \"oauth\"",
|
||
"torch; extra == \"torch\"",
|
||
"safetensors[torch]; extra == \"torch\"",
|
||
"toml; extra == \"fastai\"",
|
||
"fastai>=2.4; extra == \"fastai\"",
|
||
"fastcore>=1.3.27; extra == \"fastai\"",
|
||
"hf-xet<2.0.0,>=1.5.2; extra == \"hf-xet\"",
|
||
"mcp<2.0.0,>=1.8.0; extra == \"mcp\"",
|
||
"authlib>=1.3.2; extra == \"testing\"",
|
||
"fastapi; extra == \"testing\"",
|
||
"httpx; extra == \"testing\"",
|
||
"itsdangerous; extra == \"testing\"",
|
||
"jedi; extra == \"testing\"",
|
||
"Jinja2; extra == \"testing\"",
|
||
"pytest>=8.4.2; extra == \"testing\"",
|
||
"pytest-cov; extra == \"testing\"",
|
||
"pytest-env; extra == \"testing\"",
|
||
"pytest-xdist; extra == \"testing\"",
|
||
"pytest-vcr; extra == \"testing\"",
|
||
"pytest-asyncio; extra == \"testing\"",
|
||
"pytest-rerunfailures>=16.2; extra == \"testing\"",
|
||
"pytest-mock; extra == \"testing\"",
|
||
"urllib3<2.0; extra == \"testing\"",
|
||
"soundfile; extra == \"testing\"",
|
||
"Pillow; extra == \"testing\"",
|
||
"numpy; extra == \"testing\"",
|
||
"duckdb; extra == \"testing\"",
|
||
"fastapi; extra == \"testing\"",
|
||
"gradio>=5.0.0; extra == \"gradio\"",
|
||
"requests; extra == \"gradio\"",
|
||
"typing-extensions>=4.8.0; extra == \"typing\"",
|
||
"types-PyYAML; extra == \"typing\"",
|
||
"types-simplejson; extra == \"typing\"",
|
||
"types-toml; extra == \"typing\"",
|
||
"types-tqdm; extra == \"typing\"",
|
||
"types-urllib3; extra == \"typing\"",
|
||
"ruff>=0.9.0; extra == \"quality\"",
|
||
"mypy==1.15.0; extra == \"quality\"",
|
||
"libcst>=1.4.0; extra == \"quality\"",
|
||
"ty; extra == \"quality\"",
|
||
"authlib>=1.3.2; extra == \"all\"",
|
||
"fastapi; extra == \"all\"",
|
||
"httpx; extra == \"all\"",
|
||
"itsdangerous; extra == \"all\"",
|
||
"jedi; extra == \"all\"",
|
||
"Jinja2; extra == \"all\"",
|
||
"pytest>=8.4.2; extra == \"all\"",
|
||
"pytest-cov; extra == \"all\"",
|
||
"pytest-env; extra == \"all\"",
|
||
"pytest-xdist; extra == \"all\"",
|
||
"pytest-vcr; extra == \"all\"",
|
||
"pytest-asyncio; extra == \"all\"",
|
||
"pytest-rerunfailures>=16.2; extra == \"all\"",
|
||
"pytest-mock; extra == \"all\"",
|
||
"urllib3<2.0; extra == \"all\"",
|
||
"soundfile; extra == \"all\"",
|
||
"Pillow; extra == \"all\"",
|
||
"numpy; extra == \"all\"",
|
||
"duckdb; extra == \"all\"",
|
||
"fastapi; extra == \"all\"",
|
||
"ruff>=0.9.0; extra == \"all\"",
|
||
"mypy==1.15.0; extra == \"all\"",
|
||
"libcst>=1.4.0; extra == \"all\"",
|
||
"ty; extra == \"all\"",
|
||
"typing-extensions>=4.8.0; extra == \"all\"",
|
||
"types-PyYAML; extra == \"all\"",
|
||
"types-simplejson; extra == \"all\"",
|
||
"types-toml; extra == \"all\"",
|
||
"types-tqdm; extra == \"all\"",
|
||
"types-urllib3; extra == \"all\"",
|
||
"authlib>=1.3.2; extra == \"dev\"",
|
||
"fastapi; extra == \"dev\"",
|
||
"httpx; extra == \"dev\"",
|
||
"itsdangerous; extra == \"dev\"",
|
||
"jedi; extra == \"dev\"",
|
||
"Jinja2; extra == \"dev\"",
|
||
"pytest>=8.4.2; extra == \"dev\"",
|
||
"pytest-cov; extra == \"dev\"",
|
||
"pytest-env; extra == \"dev\"",
|
||
"pytest-xdist; extra == \"dev\"",
|
||
"pytest-vcr; extra == \"dev\"",
|
||
"pytest-asyncio; extra == \"dev\"",
|
||
"pytest-rerunfailures>=16.2; extra == \"dev\"",
|
||
"pytest-mock; extra == \"dev\"",
|
||
"urllib3<2.0; extra == \"dev\"",
|
||
"soundfile; extra == \"dev\"",
|
||
"Pillow; extra == \"dev\"",
|
||
"numpy; extra == \"dev\"",
|
||
"duckdb; extra == \"dev\"",
|
||
"fastapi; extra == \"dev\"",
|
||
"ruff>=0.9.0; extra == \"dev\"",
|
||
"mypy==1.15.0; extra == \"dev\"",
|
||
"libcst>=1.4.0; extra == \"dev\"",
|
||
"ty; extra == \"dev\"",
|
||
"typing-extensions>=4.8.0; extra == \"dev\"",
|
||
"types-PyYAML; extra == \"dev\"",
|
||
"types-simplejson; extra == \"dev\"",
|
||
"types-toml; extra == \"dev\"",
|
||
"types-tqdm; extra == \"dev\"",
|
||
"types-urllib3; extra == \"dev\""
|
||
],
|
||
"requires_python": ">=3.10.0",
|
||
"provides_extra": [
|
||
"oauth",
|
||
"torch",
|
||
"fastai",
|
||
"hf-xet",
|
||
"mcp",
|
||
"testing",
|
||
"gradio",
|
||
"typing",
|
||
"quality",
|
||
"all",
|
||
"dev"
|
||
]
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/db/f7/0a69ac6b82dbccf3f71add938a161c497952749294b8dd6dfe03a819dc40/tokenizers-0.23.2-cp310-abi3-win_amd64.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=2e96f5699d5249c9c64aa8412e044f727aae3a4098cf830f9901ec1afc361cde",
|
||
"hashes": {
|
||
"sha256": "2e96f5699d5249c9c64aa8412e044f727aae3a4098cf830f9901ec1afc361cde"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": false,
|
||
"metadata": {
|
||
"metadata_version": "2.4",
|
||
"name": "tokenizers",
|
||
"version": "0.23.2",
|
||
"description": "<p align=\"center\">\r\n <br>\r\n <img src=\"https://huggingface.co/landing/assets/tokenizers/tokenizers-logo.png\" width=\"600\"/>\r\n <br>\r\n<p>\r\n<p align=\"center\">\r\n <a href=\"https://badge.fury.io/py/tokenizers\">\r\n <img alt=\"Build\" src=\"https://badge.fury.io/py/tokenizers.svg\">\r\n </a>\r\n <a href=\"https://github.com/huggingface/tokenizers/blob/master/LICENSE\">\r\n <img alt=\"GitHub\" src=\"https://img.shields.io/github/license/huggingface/tokenizers.svg?color=blue\">\r\n </a>\r\n</p>\r\n<br>\r\n\r\n# Tokenizers\r\n\r\nProvides an implementation of today's most used tokenizers, with a focus on performance and\r\nversatility.\r\n\r\nBindings over the [Rust](https://github.com/huggingface/tokenizers/tree/master/tokenizers) implementation.\r\nIf you are interested in the High-level design, you can go check it there.\r\n\r\nOtherwise, let's dive in!\r\n\r\n## Main features:\r\n\r\n - Train new vocabularies and tokenize using 4 pre-made tokenizers (Bert WordPiece and the 3\r\n most common BPE versions).\r\n - Extremely fast (both training and tokenization), thanks to the Rust implementation. Takes\r\n less than 20 seconds to tokenize a GB of text on a server's CPU.\r\n - Easy to use, but also extremely versatile.\r\n - Designed for research and production.\r\n - Normalization comes with alignments tracking. It's always possible to get the part of the\r\n original sentence that corresponds to a given token.\r\n - Does all the pre-processing: Truncate, Pad, add the special tokens your model needs.\r\n\r\n### Installation\r\n\r\n#### With pip:\r\n\r\n```bash\r\npip install tokenizers\r\n```\r\n\r\n#### From sources:\r\n\r\nTo use this method, you need to have the Rust installed:\r\n\r\n```bash\r\n# Install with:\r\ncurl https://sh.rustup.rs -sSf | sh -s -- -y\r\nexport PATH=\"$HOME/.cargo/bin:$PATH\"\r\n```\r\n\r\nOnce Rust is installed, you can compile doing the following\r\n\r\n```bash\r\ngit clone https://github.com/huggingface/tokenizers\r\ncd tokenizers/bindings/python\r\n\r\n# Create a virtual env (you can use yours as well)\r\npython -m venv .env\r\nsource .env/bin/activate\r\n\r\n# Install `tokenizers` in the current virtual env\r\npip install -e .\r\n```\r\n\r\n### Free-threaded Python (3.14t)\r\n\r\n`tokenizers` ships dedicated wheels for the [free-threaded build of CPython](https://docs.python.org/3.14/howto/free-threading-python.html)\r\n(`python3.14t`). These wheels declare `Py_MOD_GIL_NOT_USED`, so importing\r\n`tokenizers` does **not** force the GIL back on — multi-threaded code stays\r\nGIL-free.\r\n\r\nThe full mutable API works on 3.14t — the same as on regular CPython.\r\nSetters are thread-safe: the inner tokenizer state is wrapped in a\r\n`std::sync::RwLock`, so concurrent `tokenizer.X = …` from multiple threads\r\nserialize correctly and concurrent encode operations take a read guard\r\nthat blocks writers only briefly.\r\n\r\n```python\r\nfrom tokenizers import Tokenizer\r\nfrom tokenizers.models import BPE\r\nfrom tokenizers.pre_tokenizers import Whitespace\r\nfrom tokenizers.processors import ByteLevel\r\n\r\ntok = Tokenizer(BPE())\r\ntok.pre_tokenizer = Whitespace() # ✅ thread-safe on 3.14t\r\ntok.post_processor = ByteLevel(trim_offsets=True)\r\n```\r\n\r\n**Caveat — compound mutations are not atomic.** Statements like\r\n`tokenizer.post_processor.special_tokens = X` evaluate in two steps from\r\nPython's point of view (read attribute → set attribute on the result). If\r\nanother thread swaps `tokenizer.post_processor` between those steps, the\r\nmutation lands on an orphaned component. This is the same class of race\r\nas `dict[k] = v` interleaved with `dict.clear()` — coordinate with a Python\r\nlock if you need the compound to be atomic.\r\n\r\nFor the full thread-safety analysis, see\r\n[`docs/free-threading-audit.md`](./docs/free-threading-audit.md).\r\n\r\n### Load a pretrained tokenizer from the Hub\r\n\r\n```python\r\nfrom tokenizers import Tokenizer\r\n\r\ntokenizer = Tokenizer.from_pretrained(\"bert-base-cased\")\r\n```\r\n\r\n### Using the provided Tokenizers\r\n\r\nWe provide some pre-build tokenizers to cover the most common cases. You can easily load one of\r\nthese using some `vocab.json` and `merges.txt` files:\r\n\r\n```python\r\nfrom tokenizers import CharBPETokenizer\r\n\r\n# Initialize a tokenizer\r\nvocab = \"./path/to/vocab.json\"\r\nmerges = \"./path/to/merges.txt\"\r\ntokenizer = CharBPETokenizer(vocab, merges)\r\n\r\n# And then encode:\r\nencoded = tokenizer.encode(\"I can feel the magic, can you?\")\r\nprint(encoded.ids)\r\nprint(encoded.tokens)\r\n```\r\n\r\nAnd you can train them just as simply:\r\n\r\n```python\r\nfrom tokenizers import CharBPETokenizer\r\n\r\n# Initialize a tokenizer\r\ntokenizer = CharBPETokenizer()\r\n\r\n# Then train it!\r\ntokenizer.train([ \"./path/to/files/1.txt\", \"./path/to/files/2.txt\" ])\r\n\r\n# Now, let's use it:\r\nencoded = tokenizer.encode(\"I can feel the magic, can you?\")\r\n\r\n# And finally save it somewhere\r\ntokenizer.save(\"./path/to/directory/my-bpe.tokenizer.json\")\r\n```\r\n\r\n#### Provided Tokenizers\r\n\r\n - `CharBPETokenizer`: The original BPE\r\n - `ByteLevelBPETokenizer`: The byte level version of the BPE\r\n - `SentencePieceBPETokenizer`: A BPE implementation compatible with the one used by SentencePiece\r\n - `BertWordPieceTokenizer`: The famous Bert tokenizer, using WordPiece\r\n\r\nAll of these can be used and trained as explained above!\r\n\r\n### Build your own\r\n\r\nWhenever these provided tokenizers don't give you enough freedom, you can build your own tokenizer,\r\nby putting all the different parts you need together.\r\nYou can check how we implemented the [provided tokenizers](https://github.com/huggingface/tokenizers/tree/master/bindings/python/py_src/tokenizers/implementations) and adapt them easily to your own needs.\r\n\r\n#### Building a byte-level BPE\r\n\r\nHere is an example showing how to build your own byte-level BPE by putting all the different pieces\r\ntogether, and then saving it to a single file:\r\n\r\n```python\r\nfrom tokenizers import Tokenizer, models, pre_tokenizers, decoders, trainers, processors\r\n\r\n# Initialize a tokenizer\r\ntokenizer = Tokenizer(models.BPE())\r\n\r\n# Customize pre-tokenization and decoding\r\ntokenizer.pre_tokenizer = pre_tokenizers.ByteLevel(add_prefix_space=True)\r\ntokenizer.decoder = decoders.ByteLevel()\r\ntokenizer.post_processor = processors.ByteLevel(trim_offsets=True)\r\n\r\n# And then train\r\ntrainer = trainers.BpeTrainer(\r\n vocab_size=20000,\r\n min_frequency=2,\r\n initial_alphabet=pre_tokenizers.ByteLevel.alphabet()\r\n)\r\ntokenizer.train([\r\n \"./path/to/dataset/1.txt\",\r\n \"./path/to/dataset/2.txt\",\r\n \"./path/to/dataset/3.txt\"\r\n], trainer=trainer)\r\n\r\n# And Save it\r\ntokenizer.save(\"byte-level-bpe.tokenizer.json\", pretty=True)\r\n```\r\n\r\nNow, when you want to use this tokenizer, this is as simple as:\r\n\r\n```python\r\nfrom tokenizers import Tokenizer\r\n\r\ntokenizer = Tokenizer.from_file(\"byte-level-bpe.tokenizer.json\")\r\n\r\nencoded = tokenizer.encode(\"I can feel the magic, can you?\")\r\n```\r\n\r\n### Typing support and stub generation\r\n\r\nThe compiled PyO3 extension does not expose type annotations, so editors and type checkers would otherwise see most objects as `Any`. To provide full typing support, we use a two-step stub generation process:\r\n\r\n1. **Rust introspection** (`tools/stub-gen/`): Uses `pyo3-introspection` to analyze the compiled extension and generate `.pyi` stub files\r\n2. **Python enrichment** (`stub.py`): Adds docstrings from the runtime module and generates forwarding `__init__.py` shims\r\n\r\n#### Running stub generation\r\n\r\nThe easiest way to regenerate stubs is via `make style`:\r\n\r\n```bash\r\ncd bindings/python\r\nmake style\r\n```\r\n\r\nThis will:\r\n1. Build the extension with `maturin develop --release`\r\n2. Run introspection to generate `.pyi` files\r\n3. Enrich stubs with docstrings via `stub.py`\r\n4. Format with `ruff`\r\n\r\n#### Running manually\r\n\r\nTo run the stub generator directly:\r\n\r\n```bash\r\ncd bindings/python\r\ncargo run --manifest-path tools/stub-gen/Cargo.toml\r\npython stub.py\r\n```\r\n\r\nThe stub generator automatically:\r\n- Builds the extension using maturin\r\n- Copies the built `.so` to the project root for introspection\r\n- Detects and sets `PYTHONHOME` for embedded Python (handles uv/venv environments)\r\n- Generates stubs to `py_src/tokenizers/`\r\n\r\n#### Troubleshooting\r\n\r\nIf you encounter Python initialization errors, you can manually set `PYTHONHOME`:\r\n\r\n```bash\r\nexport PYTHONHOME=$(python3 -c 'import sys; print(sys.base_prefix)')\r\ncargo run --manifest-path tools/stub-gen/Cargo.toml\r\n```\r\n\n",
|
||
"description_content_type": "text/markdown; charset=UTF-8; variant=GFM",
|
||
"keywords": [
|
||
"NLP",
|
||
"tokenizer",
|
||
"BPE",
|
||
"transformer",
|
||
"deep learning"
|
||
],
|
||
"author_email": "Nicolas Patry <patry.nicolas@protonmail.com>, Anthony Moi <anthony@huggingface.co>",
|
||
"classifier": [
|
||
"Development Status :: 5 - Production/Stable",
|
||
"Intended Audience :: Developers",
|
||
"Intended Audience :: Education",
|
||
"Intended Audience :: Science/Research",
|
||
"License :: OSI Approved :: Apache Software License",
|
||
"Operating System :: OS Independent",
|
||
"Programming Language :: Python :: 3",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: 3.14",
|
||
"Programming Language :: Python :: 3 :: Only",
|
||
"Topic :: Scientific/Engineering :: Artificial Intelligence"
|
||
],
|
||
"requires_dist": [
|
||
"huggingface-hub>=0.16.4,<2.0",
|
||
"tokenizers[testing] ; extra == 'dev'",
|
||
"sphinx ; extra == 'docs'",
|
||
"sphinx-rtd-theme ; extra == 'docs'",
|
||
"setuptools-rust ; extra == 'docs'",
|
||
"pytest ; extra == 'testing'",
|
||
"pytest-asyncio ; extra == 'testing'",
|
||
"requests ; extra == 'testing'",
|
||
"numpy ; extra == 'testing'",
|
||
"datasets ; extra == 'testing'",
|
||
"ruff ; extra == 'testing'",
|
||
"ty ; extra == 'testing'"
|
||
],
|
||
"requires_python": ">=3.10",
|
||
"project_url": [
|
||
"Homepage, https://github.com/huggingface/tokenizers",
|
||
"Source, https://github.com/huggingface/tokenizers"
|
||
],
|
||
"provides_extra": [
|
||
"dev",
|
||
"docs",
|
||
"testing"
|
||
]
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/01/fb/0e4489505dac46a0487ba925ea4b046612f1a40f432d63a000421de78549/filelock-4.0.0-py3-none-any.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=a850aa9ec2acba8db9ca2e9fcf8a327fbc2f85e432725715b37bd76b7dd1f798",
|
||
"hashes": {
|
||
"sha256": "a850aa9ec2acba8db9ca2e9fcf8a327fbc2f85e432725715b37bd76b7dd1f798"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": false,
|
||
"metadata": {
|
||
"metadata_version": "2.5",
|
||
"name": "filelock",
|
||
"version": "4.0.0",
|
||
"summary": "A platform independent file lock.",
|
||
"description": "# filelock\n\n[](https://pypi.org/project/filelock/)\n[](https://pypi.org/project/filelock/)\n[](https://py-filelock.readthedocs.io/en/latest/?badge=latest)\n[](https://pepy.tech/project/filelock)\n[](https://github.com/tox-dev/py-filelock/actions/workflows/check.yaml)\n\nFor more information checkout the [official documentation](https://py-filelock.readthedocs.io/en/latest/index.html).\n",
|
||
"description_content_type": "text/markdown",
|
||
"keywords": [
|
||
"application",
|
||
"cache",
|
||
"directory",
|
||
"log",
|
||
"user"
|
||
],
|
||
"maintainer_email": "Bernát Gábor <gaborjbernat@gmail.com>",
|
||
"license_expression": "MIT",
|
||
"license_file": [
|
||
"LICENSE"
|
||
],
|
||
"classifier": [
|
||
"Development Status :: 5 - Production/Stable",
|
||
"Intended Audience :: Developers",
|
||
"License :: OSI Approved :: MIT License",
|
||
"Operating System :: OS Independent",
|
||
"Programming Language :: Python",
|
||
"Programming Language :: Python :: 3 :: Only",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: 3.14",
|
||
"Programming Language :: Python :: 3.15",
|
||
"Topic :: Internet",
|
||
"Topic :: Software Development :: Libraries",
|
||
"Topic :: System"
|
||
],
|
||
"requires_python": ">=3.10",
|
||
"project_url": [
|
||
"Documentation, https://py-filelock.readthedocs.io",
|
||
"Homepage, https://github.com/tox-dev/py-filelock",
|
||
"Source, https://github.com/tox-dev/py-filelock",
|
||
"Tracker, https://github.com/tox-dev/py-filelock/issues"
|
||
]
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/fd/3c/6a2bf344106328fd04963664a60b9bb6496fc25df8e962fcdc1367285fb9/fsspec-2026.7.0-py3-none-any.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=b57ddbafedfaef7018c1ecab32aa200a9d7ca26b77965f64e48b70061249d279",
|
||
"hashes": {
|
||
"sha256": "b57ddbafedfaef7018c1ecab32aa200a9d7ca26b77965f64e48b70061249d279"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": false,
|
||
"metadata": {
|
||
"metadata_version": "2.4",
|
||
"name": "fsspec",
|
||
"version": "2026.7.0",
|
||
"summary": "File-system specification",
|
||
"description": "# filesystem_spec\n\n[](https://pypi.python.org/pypi/fsspec/)\n[](https://anaconda.org/conda-forge/fsspec)\n\n[](https://filesystem-spec.readthedocs.io/en/latest/?badge=latest)\n\nA specification for pythonic filesystems.\n\n## Install\n\n```bash\npip install fsspec\n```\n\nwould install the base fsspec. Various optionally supported features might require specification of custom\nextra require, e.g. `pip install fsspec[ssh]` will install dependencies for `ssh` backends support.\nUse `pip install fsspec[full]` for installation of all known extra dependencies.\n\nUp-to-date package also provided through conda-forge distribution:\n\n```bash\nconda install -c conda-forge fsspec\n```\n\n\n## Purpose\n\nTo produce a template or specification for a file-system interface, that specific implementations should follow,\nso that applications making use of them can rely on a common behaviour and not have to worry about the specific\ninternal implementation decisions with any given backend. Many such implementations are included in this package,\nor in sister projects such as `s3fs` and `gcsfs`.\n\nIn addition, if this is well-designed, then additional functionality, such as a key-value store or FUSE\nmounting of the file-system implementation may be available for all implementations \"for free\".\n\n## Documentation\n\nPlease refer to [RTD](https://filesystem-spec.readthedocs.io/en/latest/?badge=latest)\n\n## Develop\n\nfsspec uses GitHub Actions for CI. Environment files can be found\nin the \"ci/\" directory. Note that the main environment is called \"py38\",\nbut it is expected that the version of python installed be adjustable at\nCI runtime. For local use, pick a version suitable for you.\n\n```bash\n# For a new environment (mamba / conda).\nmamba create -n fsspec -c conda-forge python=3.10 -y\nconda activate fsspec\n\n# Standard dev install with docs and tests.\npip install -e \".[dev,doc,test]\"\n\n# Full tests except for downstream\npip install s3fs\npip uninstall s3fs\npip install -e .[dev,doc,test_full]\npip install s3fs --no-deps\npytest -v\n\n# Downstream tests.\nsh install_s3fs.sh\n# Windows powershell.\ninstall_s3fs.sh\n```\n\n### Testing\n\nTests can be run in the dev environment, if activated, via ``pytest fsspec``.\n\nThe full fsspec suite requires a system-level docker, docker-compose, and fuse\ninstallation. If only making changes to one backend implementation, it is\nnot generally necessary to run all tests locally.\n\nIt is expected that contributors ensure that any change to fsspec does not\ncause issues or regressions for either other fsspec-related packages such\nas gcsfs and s3fs, nor for downstream users of fsspec. The \"downstream\" CI\nrun and corresponding environment file run a set of tests from the dask\ntest suite, and very minimal tests against pandas and zarr from the\ntest_downstream.py module in this repo.\n\n### Code Formatting\n\nfsspec uses [Black](https://black.readthedocs.io/en/stable) to ensure\na consistent code format throughout the project.\nRun ``black fsspec`` from the root of the filesystem_spec repository to\nauto-format your code. Additionally, many editors have plugins that will apply\n``black`` as you edit files. ``black`` is included in the ``tox`` environments.\n\nOptionally, you may wish to setup [pre-commit hooks](https://pre-commit.com) to\nautomatically run ``black`` when you make a git commit.\nRun ``pre-commit install --install-hooks`` from the root of the\nfilesystem_spec repository to setup pre-commit hooks. ``black`` will now be run\nbefore you commit, reformatting any changed files. You can format without\ncommitting via ``pre-commit run`` or skip these checks with ``git commit\n--no-verify``.\n\n## Support\n\nWork on this repository is supported in part by:\n\n\"Anaconda, Inc. - Advancing AI through open source.\"\n\n<a href=\"https://anaconda.com/\"><img src=\"https://camo.githubusercontent.com/b8555ef2222598ed37ce38ac86955febbd25de7619931bb7dd3c58432181d3b6/68747470733a2f2f626565776172652e6f72672f636f6d6d756e6974792f6d656d626572732f616e61636f6e64612f616e61636f6e64612d6c617267652e706e67\" alt=\"anaconda logo\" width=\"40%\"/></a>\n",
|
||
"description_content_type": "text/markdown",
|
||
"keywords": [
|
||
"file"
|
||
],
|
||
"maintainer_email": "Martin Durant <mdurant@anaconda.com>",
|
||
"license_expression": "BSD-3-Clause",
|
||
"license_file": [
|
||
"LICENSE"
|
||
],
|
||
"classifier": [
|
||
"Development Status :: 4 - Beta",
|
||
"Intended Audience :: Developers",
|
||
"Operating System :: OS Independent",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: 3.14"
|
||
],
|
||
"requires_dist": [
|
||
"adlfs; extra == 'abfs'",
|
||
"adlfs; extra == 'adl'",
|
||
"pyarrow>=1; extra == 'arrow'",
|
||
"dask; extra == 'dask'",
|
||
"distributed; extra == 'dask'",
|
||
"pre-commit; extra == 'dev'",
|
||
"ruff>=0.5; extra == 'dev'",
|
||
"numpydoc; extra == 'doc'",
|
||
"sphinx; extra == 'doc'",
|
||
"sphinx-design; extra == 'doc'",
|
||
"sphinx-rtd-theme; extra == 'doc'",
|
||
"yarl; extra == 'doc'",
|
||
"dropbox; extra == 'dropbox'",
|
||
"dropboxdrivefs; extra == 'dropbox'",
|
||
"requests; extra == 'dropbox'",
|
||
"adlfs; extra == 'full'",
|
||
"aiohttp!=4.0.0a0,!=4.0.0a1; extra == 'full'",
|
||
"dask; extra == 'full'",
|
||
"distributed; extra == 'full'",
|
||
"dropbox; extra == 'full'",
|
||
"dropboxdrivefs; extra == 'full'",
|
||
"fusepy; extra == 'full'",
|
||
"gcsfs>=2026.4.0; extra == 'full'",
|
||
"libarchive-c; extra == 'full'",
|
||
"ocifs; extra == 'full'",
|
||
"panel; extra == 'full'",
|
||
"paramiko; extra == 'full'",
|
||
"pyarrow>=1; extra == 'full'",
|
||
"pygit2; extra == 'full'",
|
||
"requests; extra == 'full'",
|
||
"s3fs>=2026.6.0; extra == 'full'",
|
||
"smbprotocol; extra == 'full'",
|
||
"tqdm; extra == 'full'",
|
||
"fusepy; extra == 'fuse'",
|
||
"gcsfs>=2026.4.0; extra == 'gcs'",
|
||
"pygit2; extra == 'git'",
|
||
"requests; extra == 'github'",
|
||
"gcsfs>=2026.4.0; extra == 'gs'",
|
||
"panel; extra == 'gui'",
|
||
"pyarrow>=1; extra == 'hdfs'",
|
||
"aiohttp!=4.0.0a0,!=4.0.0a1; extra == 'http'",
|
||
"libarchive-c; extra == 'libarchive'",
|
||
"ocifs; extra == 'oci'",
|
||
"s3fs>=2026.6.0; extra == 's3'",
|
||
"paramiko; extra == 'sftp'",
|
||
"smbprotocol; extra == 'smb'",
|
||
"paramiko; extra == 'ssh'",
|
||
"aiohttp!=4.0.0a0,!=4.0.0a1; extra == 'test'",
|
||
"numpy; extra == 'test'",
|
||
"pytest; extra == 'test'",
|
||
"pytest-asyncio!=0.22.0; extra == 'test'",
|
||
"pytest-benchmark; extra == 'test'",
|
||
"pytest-cov; extra == 'test'",
|
||
"pytest-mock; extra == 'test'",
|
||
"pytest-recording; extra == 'test'",
|
||
"pytest-rerunfailures; extra == 'test'",
|
||
"requests; extra == 'test'",
|
||
"aiobotocore<3.0.0,>=2.5.4; extra == 'test-downstream'",
|
||
"dask[dataframe,test]; extra == 'test-downstream'",
|
||
"moto[server]<5,>4; extra == 'test-downstream'",
|
||
"pytest-timeout; extra == 'test-downstream'",
|
||
"xarray; extra == 'test-downstream'",
|
||
"adlfs; extra == 'test-full'",
|
||
"aiohttp!=4.0.0a0,!=4.0.0a1; extra == 'test-full'",
|
||
"backports-zstd; (python_version < '3.14') and extra == 'test-full'",
|
||
"cloudpickle; extra == 'test-full'",
|
||
"dask; extra == 'test-full'",
|
||
"distributed; extra == 'test-full'",
|
||
"dropbox; extra == 'test-full'",
|
||
"dropboxdrivefs; extra == 'test-full'",
|
||
"fastparquet; extra == 'test-full'",
|
||
"fusepy; extra == 'test-full'",
|
||
"gcsfs>=2026.4.0; extra == 'test-full'",
|
||
"jinja2; extra == 'test-full'",
|
||
"kerchunk; extra == 'test-full'",
|
||
"libarchive-c; extra == 'test-full'",
|
||
"lz4; extra == 'test-full'",
|
||
"notebook; extra == 'test-full'",
|
||
"numpy; extra == 'test-full'",
|
||
"ocifs; extra == 'test-full'",
|
||
"pandas<3.0.0; extra == 'test-full'",
|
||
"panel; extra == 'test-full'",
|
||
"paramiko; extra == 'test-full'",
|
||
"pyarrow>=1; extra == 'test-full'",
|
||
"pyftpdlib; extra == 'test-full'",
|
||
"pygit2; extra == 'test-full'",
|
||
"pytest; extra == 'test-full'",
|
||
"pytest-asyncio!=0.22.0; extra == 'test-full'",
|
||
"pytest-benchmark; extra == 'test-full'",
|
||
"pytest-cov; extra == 'test-full'",
|
||
"pytest-mock; extra == 'test-full'",
|
||
"pytest-recording; extra == 'test-full'",
|
||
"pytest-rerunfailures; extra == 'test-full'",
|
||
"python-snappy; extra == 'test-full'",
|
||
"requests; extra == 'test-full'",
|
||
"s3fs>=2026.6.0; extra == 'test-full'",
|
||
"smbprotocol; extra == 'test-full'",
|
||
"tqdm; extra == 'test-full'",
|
||
"urllib3; extra == 'test-full'",
|
||
"zarr<3.2.0; extra == 'test-full'",
|
||
"zstandard; (python_version < '3.14') and extra == 'test-full'",
|
||
"tqdm; extra == 'tqdm'"
|
||
],
|
||
"requires_python": ">=3.10",
|
||
"project_url": [
|
||
"Changelog, https://filesystem-spec.readthedocs.io/en/latest/changelog.html",
|
||
"Documentation, https://filesystem-spec.readthedocs.io/en/latest/",
|
||
"Homepage, https://github.com/fsspec/filesystem_spec"
|
||
],
|
||
"provides_extra": [
|
||
"abfs",
|
||
"adl",
|
||
"arrow",
|
||
"dask",
|
||
"dev",
|
||
"doc",
|
||
"dropbox",
|
||
"entrypoints",
|
||
"full",
|
||
"fuse",
|
||
"gcs",
|
||
"git",
|
||
"github",
|
||
"gs",
|
||
"gui",
|
||
"hdfs",
|
||
"http",
|
||
"libarchive",
|
||
"oci",
|
||
"s3",
|
||
"sftp",
|
||
"smb",
|
||
"ssh",
|
||
"test",
|
||
"test-downstream",
|
||
"test-full",
|
||
"tqdm"
|
||
]
|
||
}
|
||
},
|
||
{
|
||
"download_info": {
|
||
"url": "https://files.pythonhosted.org/packages/98/b7/8c59a66d15205024662f1d66968136f13893f96df1ddc5087e2e281fc95f/hf_xet-1.6.0-cp38-abi3-win_amd64.whl",
|
||
"archive_info": {
|
||
"hash": "sha256=fb4fadde1b2b70bf4c0c14a6dccbe7194b1c28947fefd5bbe3fed9d940676c3b",
|
||
"hashes": {
|
||
"sha256": "fb4fadde1b2b70bf4c0c14a6dccbe7194b1c28947fefd5bbe3fed9d940676c3b"
|
||
}
|
||
}
|
||
},
|
||
"is_direct": false,
|
||
"is_yanked": false,
|
||
"requested": false,
|
||
"metadata": {
|
||
"metadata_version": "2.4",
|
||
"name": "hf-xet",
|
||
"version": "1.6.0",
|
||
"summary": "Fast transfer of large files with the Hugging Face Hub.",
|
||
"description": "<!---\r\nCopyright 2024 The HuggingFace Team. All rights reserved.\r\n\r\nLicensed under the Apache License, Version 2.0 (the \"License\");\r\nyou may not use this file except in compliance with the License.\r\nYou may obtain a copy of the License at\r\n\r\n http://www.apache.org/licenses/LICENSE-2.0\r\n\r\nUnless required by applicable law or agreed to in writing, software\r\ndistributed under the License is distributed on an \"AS IS\" BASIS,\r\nWITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\r\nSee the License for the specific language governing permissions and\r\nlimitations under the License.\r\n-->\r\n<p align=\"center\">\r\n <a href=\"https://github.com/huggingface/xet-core/blob/main/LICENSE\"><img alt=\"License\" src=\"https://img.shields.io/github/license/huggingface/xet-core.svg?color=blue\"></a>\r\n <a href=\"https://github.com/huggingface/xet-core/releases\"><img alt=\"GitHub release\" src=\"https://img.shields.io/github/release/huggingface/xet-core.svg\"></a>\r\n <a href=\"https://github.com/huggingface/xet-core/blob/main/CODE_OF_CONDUCT.md\"><img alt=\"Contributor Covenant\" src=\"https://img.shields.io/badge/Contributor%20Covenant-v2.0%20adopted-ff69b4.svg\"></a>\r\n</p>\r\n\r\n<h3 align=\"center\">\r\n <p>🤗 hf-xet - xet client tech, used in <a target=\"_blank\" href=\"https://github.com/huggingface/huggingface_hub/\">huggingface_hub</a></p>\r\n</h3>\r\n\r\n## Welcome\r\n\r\n`hf-xet` enables `huggingface_hub` to utilize xet storage for uploading and downloading to HF Hub. Xet storage provides chunk-based deduplication, efficient storage/retrieval with local disk caching, and backwards compatibility with Git LFS. This library is not meant to be used directly, and is instead intended to be used from [huggingface_hub](https://pypi.org/project/huggingface-hub).\r\n\r\n## Key features\r\n\r\n♻ **chunk-based deduplication implementation**: avoid transferring and storing chunks that are shared across binary files (models, datasets, etc).\r\n\r\n🤗 **Python bindings**: bindings for [huggingface_hub](https://github.com/huggingface/huggingface_hub/) package.\r\n\r\n↔ **network communications**: concurrent communication to HF Hub Xet backend services (CAS).\r\n\r\n🔖 **local disk caching**: chunk-based cache that sits alongside the existing [huggingface_hub disk cache](https://huggingface.co/docs/huggingface_hub/guides/manage-cache).\r\n\r\n## Installation\r\n\r\nInstall the `hf_xet` package with [pip](https://pypi.org/project/hf-xet/):\r\n\r\n```bash\r\npip install hf_xet\r\n```\r\n\r\n## Quick Start\r\n\r\n`hf_xet` is not intended to be run independently as it is expected to be used from `huggingface_hub`, so to get started with `huggingface_hub` check out the documentation [here](\"https://hf.co/docs/huggingface_hub\").\r\n\r\n## Contributions (feature requests, bugs, etc.) are encouraged & appreciated 💙💚💛💜🧡❤️\r\n\r\nPlease join us in making hf-xet better. We value everyone's contributions. Code is not the only way to help. Answering questions, helping each other, improving documentation, filing issues all help immensely. If you are interested in contributing (please do!), check out the [contribution guide](https://github.com/huggingface/xet-core/blob/main/CONTRIBUTING.md) for this repository.\n",
|
||
"description_content_type": "text/markdown; charset=UTF-8; variant=GFM",
|
||
"maintainer_email": "Rajat Arya <rajat@rajatarya.com>, Jared Sulzdorf <j.sulzdorf@gmail.com>, Di Xiao <di@huggingface.co>, Assaf Vayner <assaf@huggingface.co>, Hoyt Koepke <hoytak@gmail.com>",
|
||
"license_expression": "Apache-2.0",
|
||
"license_file": [
|
||
"LICENSE"
|
||
],
|
||
"classifier": [
|
||
"Development Status :: 5 - Production/Stable",
|
||
"License :: OSI Approved :: Apache Software License",
|
||
"Programming Language :: Rust",
|
||
"Programming Language :: Python :: Implementation :: CPython",
|
||
"Programming Language :: Python :: Implementation :: PyPy",
|
||
"Programming Language :: Python :: 3",
|
||
"Programming Language :: Python :: 3 :: Only",
|
||
"Programming Language :: Python :: 3.8",
|
||
"Programming Language :: Python :: 3.9",
|
||
"Programming Language :: Python :: 3.10",
|
||
"Programming Language :: Python :: 3.11",
|
||
"Programming Language :: Python :: 3.12",
|
||
"Programming Language :: Python :: 3.13",
|
||
"Programming Language :: Python :: 3.14",
|
||
"Programming Language :: Python :: Free Threading",
|
||
"Programming Language :: Python :: Free Threading :: 2 - Beta",
|
||
"Topic :: Scientific/Engineering :: Artificial Intelligence"
|
||
],
|
||
"requires_dist": [
|
||
"pytest ; extra == 'tests'"
|
||
],
|
||
"requires_python": ">=3.8",
|
||
"project_url": [
|
||
"Documentation, https://huggingface.co/docs/hub/xet/index",
|
||
"Homepage, https://github.com/huggingface/xet-core",
|
||
"Issues, https://github.com/huggingface/xet-core/issues",
|
||
"Repository, https://github.com/huggingface/xet-core.git"
|
||
],
|
||
"provides_extra": [
|
||
"tests"
|
||
]
|
||
}
|
||
}
|
||
],
|
||
"environment": {
|
||
"implementation_name": "cpython",
|
||
"implementation_version": "3.12.14",
|
||
"os_name": "nt",
|
||
"platform_machine": "AMD64",
|
||
"platform_release": "11",
|
||
"platform_system": "Windows",
|
||
"platform_version": "10.0.26200",
|
||
"python_full_version": "3.12.14",
|
||
"platform_python_implementation": "CPython",
|
||
"python_version": "3.12",
|
||
"sys_platform": "win32"
|
||
}
|
||
} |