AI Tools.

Search

image text to text models

293 models · ranked by HuggingFace downloads

Qwen3.6-35B-A3B-FP8

FP8-quantized version of Qwen3.6-35B-A3B for deployment on hardware with FP8 support (H100/H200). Reduces memory footprint and inference latency compared to BF16 with minimal quality degradation on most benchmarks.

12,944,800 ↓ · 359 ♡

Qwen3.5-9B

Qwen3.5-9B is a 9-billion-parameter instruction-tuned vision-language model from Alibaba Cloud's Qwen3.5 series, fine-tuned from Qwen3.5-9B-Base for multimodal conversational tasks. It accepts image and text inputs for visual reasoning, document understanding, and grounded question answering. Apache 2.0 licensed.

12,507,912 ↓ · 1,876 ♡

Qwen3.6-27B-FP8

FP8-quantized version of Qwen 3.6 27B for H100/H200 serving. Reduces memory from ~54GB (BF16) to approximately 27GB while maintaining near-BF16 quality on most benchmarks for a dense multimodal model.

8,373,508 ↓ · 349 ♡

gemma-4-31B-it

Gemma 4-31B-IT is Google DeepMind's 31-billion-parameter instruction-tuned vision-language model from the Gemma 4 family, supporting both image and text inputs. It offers strong multimodal reasoning at open-weight scale, with Apache 2.0 licensing making it directly deployable for commercial applications. Part of the gemma4 architecture with improvements over Gemma 2.

8,354,717 ↓ · 3,664 ♡

Qwen2.5-VL-7B-Instruct

Qwen2.5-VL-7B-Instruct is Alibaba Cloud's 7-billion-parameter vision-language model from the Qwen2.5-VL series, accepting image and video inputs alongside text for visual question answering, document understanding, and grounding tasks. It supports multiple image resolutions dynamically and shows improved OCR and document reasoning compared to the earlier Qwen-VL series. Apache 2.0 licensed.

8,177,889 ↓ · 1,688 ♡

gemma-4-26B-A4B-it

Gemma 4-26B-A4B-IT is Google DeepMind's 26-billion-total-parameter MoE (Mixture-of-Experts) vision-language model, with approximately 4 billion active parameters per token. The MoE design means it achieves 26B parameter quality while activating only ~4B per forward pass, reducing per-token compute relative to a dense 26B model. Apache 2.0 licensed.

8,087,656 ↓ · 1,452 ♡

Qwen3-VL-8B-Instruct

Qwen3-VL-8B-Instruct is Alibaba Cloud's 8-billion-parameter vision-language model from the Qwen3-VL series, extending the VL line with improved visual reasoning and document understanding. It targets mid-tier server GPU deployment where 2B VLMs are insufficient and 30B+ is impractical. Apache 2.0 licensed.

7,618,848 ↓ · 1,072 ♡

Qwen3.5-4B

Qwen3.5-4B is Alibaba Cloud's 4-billion-parameter instruction-tuned vision-language model from the Qwen3.5 series, fine-tuned from Qwen3.5-4B-Base for multimodal conversational tasks. It handles image and text inputs at a scale deployable on consumer GPUs with 8-12GB VRAM. Apache 2.0 licensed.

7,438,709 ↓ · 862 ♡

Qwen3.6-27B

Qwen3.6-27B is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps Qwen3.6-27B unrestricted for commercial reuse. Weighing in near 27000M parameters, Qwen3.6-27B trades some ceiling for cheaper, faster inference. Treat Qwen3.6-27B's published metrics as a starting point and validate against your workload.

5,745,593 ↓ · 2,282 ♡

Qwen3.6-35B-A3B

Qwen 3.6 is a Mixture-of-Experts model with 35B total parameters but only 3B active per token, giving MoE inference efficiency at near-35B capacity. It handles image and text inputs and is competitive with dense 14–20B models on standard benchmarks.

4,996,471 ↓ · 2,748 ♡

Qwen3.8-27B-FP8

Official FP8-quantized checkpoint of Qwen3.8-27B targeting NVIDIA H100/H200 GPUs. Cuts VRAM from roughly 54GB at BF16 to approximately 27GB while retaining near-BF16 generation quality for text and image tasks. Requires vLLM or a Transformers backend with FP8 support.

4,606,343 ↓ · 720 ♡

Qwen2.5-VL-3B-Instruct

Qwen2.5-VL-3B-Instruct is Alibaba's 3B parameter vision-language model from the Qwen2.5-VL series, supporting image and video frame understanding alongside text instruction-following. It targets edge and mobile deployment where 7B+ VL models are too memory-intensive, while maintaining reasonable accuracy on OCR, chart reading, and visual QA. Instruction-tuned for conversational use.

4,399,514 ↓ · 690 ♡

Qwen3.8-27B

Alibaba's 27B multimodal model that accepts images alongside text prompts in a single BF16 checkpoint. Competitive with similarly sized models on vision benchmarks and code tasks. Apache 2.0 licensed with Azure deploy integration.

4,028,839 ↓ · 13,275 ♡

Qwen3-VL-4B-Instruct

Qwen3-VL 4B is Alibaba's compact vision-language instruction model supporting image and video understanding at 4B scale. It targets use cases where Qwen2-VL-7B quality is acceptable but deployment must fit tighter memory constraints.

3,820,376 ↓ · 453 ♡

Qwen3.6-27B-NVFP4

Unsloth's NVFP4 (NVIDIA FP4) quantization of Qwen3.6-27B, targeting inference on H100/H200 GPUs with FP4 hardware support. FP4 enables significant throughput gains over BF16 on Ada Lovelace and Hopper-architecture GPUs that support native FP4 compute.

3,432,631 ↓ · 277 ♡

Unlimited-OCR

Baidu's Unlimited-OCR is a vision-language model targeting text recognition across multiple scripts, layouts, and document types. Accompanies a preprint (arXiv:2606.23050) and ships with published eval results on standard OCR benchmarks.

2,957,934 ↓ · 4,147 ♡

Qwen3.5-2B

Qwen3.5-2B targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. It is a fine-tune of qwen3.5-2b-base, inheriting that base model's general competence. Permissive Apache 2.0 terms let Qwen3.5-2B go straight into commercial pipelines. Qwen3.5-2B is community-maintained, so track upstream changes and pin a known-good revision.

2,868,560 ↓ · 375 ♡

chandra-ocr-2

chandra-ocr-2 is an open-weight checkpoint for vision-language understanding, distributed on the HuggingFace Hub. chandra-ocr-2 is subject to OpenRAIL terms, so confirm licensing before commercial use. Evaluate chandra-ocr-2 on your own data before trusting it in production.

2,739,661 ↓ · 481 ♡

Kimi-K3

Kimi-K3 is Moonshot AI's large-scale multimodal model designed for image-text understanding and reasoning tasks. With over 9,000 community likes it is among the most widely adopted recent open-weight multimodal releases. Covers both visual comprehension and language reasoning in a single model. Non-standard license — check Moonshot AI's terms.

2,701,014 ↓ · 11,083 ♡

Qwen3-VL-2B-Instruct

Qwen3-VL-2B-Instruct is a 2-billion-parameter vision-language model from Alibaba Cloud that jointly processes images and text for visual question answering, captioning, and document understanding. Its 2B scale positions it as one of the smaller instruction-tuned VLMs capable of zero-shot visual reasoning. Apache 2.0 licensed.

2,678,334 ↓ · 452 ♡

Florence-2-base

Built for vision-language understanding, Florence-2-base is a florence-based model with publicly available weights. Florence-2-base is MIT-licensed, clearing it for closed-source and paid products. Florence-2-base ships without a hosted SLA, so budget for self-managed deployment and monitoring.

2,674,841 ↓ · 396 ♡

Gemma-4-E4B-Uncensored-HauhauCS-Aggressive

Gemma-4-E4B-Uncensored-HauhauCS-Aggressive is a gemma-based open-weight model aimed at vision-language understanding. Training spans multiple languages, so Gemma-4-E4B-Uncensored-HauhauCS-Aggressive covers cross-lingual vision-language understanding from one checkpoint. Because Gemma-4-E4B-Uncensored-HauhauCS-Aggressive uses Gemma, vet the conditions against your deployment plan. Gemma-4-E4B-Uncensored-HauhauCS-Aggressive ships without a hosted SLA, so budget for self-managed deployment and monitoring.

2,631,763 ↓ · 1,062 ♡

Qwen3.5-27B

Qwen 3.5 27B is a dense image-text-to-text model from Alibaba, positioned between the 14B and 72B variants for users who need more capacity than 14B but can't serve 72B. It handles both vision and language instructions.

2,501,686 ↓ · 1,036 ♡

gemma-4-26B-A4B-it-AWQ-4bit

An AWQ 4-bit quantized version of Gemma 4's 26B MoE model (4B active parameters), reducing the memory footprint for local deployment on consumer hardware. Community-produced quantization targeting llama.cpp and vLLM compatibility.

2,501,205 ↓ · 94 ♡

Qwen3.5-35B-A3B

Qwen3.5-35B-A3B is a 35B total parameter mixture-of-experts multimodal model from Alibaba, with approximately 3B active parameters per token during inference. It combines vision and language understanding for image captioning, visual QA, and document analysis tasks at lower compute cost than a dense 35B model. Apache 2.0 licensed.

2,473,014 ↓ · 1,496 ♡

Qwen3.5-0.8B

Qwen 3.5 0.8B is Alibaba's smallest production language model in the 3.5 series, designed for on-device and edge inference. Despite its size, it supports the same instruction format as larger Qwen models and is suitable for simple classification, extraction, and short-form generation.

2,444,153 ↓ · 679 ♡

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

A heavily merged and abliterated GGUF release from Bartowski combining Qwen3.6-27B with multiple community fine-tunes (Fable, Fusion-711, Heretic, NEO, MAX, MTP). The long name reflects the layered merge and abliteration applied to create a multi-capability uncensored conversational model.

2,422,859 ↓ · 2,277 ♡

DeepSeek-OCR

DeepSeek OCR is a vision-language model from DeepSeek optimized specifically for optical character recognition from natural scene and document images. It aims to handle mixed layouts, multi-language text, and complex typographic scenarios.

2,352,441 ↓ · 3,348 ♡

Qwen3.6-35B-A3B-NVFP4

Unsloth's NVFP4 (4-bit FP) quantization of the Qwen3.6 35B A3B mixture-of-experts model using the compressed-tensors format. The NVFP4 format targets NVIDIA's latest Blackwell GPUs with dedicated FP4 tensor cores for ultra-low-latency inference at 8-bit effective compression.

2,304,785 ↓ · 116 ♡

Qwen3-VL-8B-Instruct-FP8

Built for vision-language understanding, Qwen3-VL-8B-Instruct-FP8 is a qwen3-based model with publicly available weights. Qwen3-VL-8B-Instruct-FP8 is Apache 2.0-licensed, clearing it for closed-source and paid products. At about 8000M parameters, Qwen3-VL-8B-Instruct-FP8 sits in the large tier, which sets its memory and latency budget. Read Qwen3-VL-8B-Instruct-FP8's card for hardware requirements and licensing fine print before deploying.

2,252,996 ↓ · 79 ♡

Qwen3.8-27B-MLX-4bit

4-bit MLX quantization of Qwen3.8-27B for Apple Silicon inference via the MLX framework. Uses native Apple Metal for computation, enabling the 27B multimodal model on unified memory Mac systems with 32GB+ RAM. Apache 2.0 licensed.

2,238,624 ↓ · 30 ♡

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is a frontier-scale checkpoint for vision-language understanding, distributed on the HuggingFace Hub. Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is multilingual by design rather than English-only. The Apache 2.0 license keeps Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive unrestricted for commercial reuse. Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is community-maintained, so track upstream changes and pin a known-good revision.

2,236,537 ↓ · 3,545 ♡

Qwen2.5-VL-32B-Instruct-AWQ

Alibaba's AWQ 4-bit quantization of Qwen2.5-VL-32B-Instruct, a 32B vision-language model supporting multimodal chat. AWQ preserves critical weight channels to minimize accuracy loss compared to naive 4-bit schemes.

2,171,210 ↓ · 64 ♡

llava-1.5-7b-hf

LLaVA 1.5 7B connects a CLIP ViT-L/14@336 vision encoder to Vicuna 7B via a simple MLP projection. It was a state-of-the-art open multimodal model at release and remains widely used as a baseline for vision-language research.

2,144,298 ↓ · 372 ♡

Qwen3.8-27B-MLX-8bit

An 8-bit MLX quantization of Qwen3.8's 27B parameter multimodal model, packaged by LM Studio for Apple Silicon. Runs natively via the MLX framework, bypassing GPU memory constraints typical of larger models on M-series Macs. Based on Qwen/Qwen3.8-27B, it retains image understanding alongside text generation.

2,099,338 ↓ · 15 ♡

Qwen3.6-27B-AWQ-INT4

Qwen3.6-27B-AWQ-INT4 is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps Qwen3.6-27B-AWQ-INT4 unrestricted for commercial reuse. Prebuilt AWQ/INT4 weights make local and edge inference of Qwen3.6-27B-AWQ-INT4 straightforward. Qwen3.6-27B-AWQ-INT4 is community-maintained, so track upstream changes and pin a known-good revision.

2,068,591 ↓ · 110 ♡

Qwen3.8-27B-MLX-6bit

A 6-bit MLX quantization of Qwen3.8-27B packaged for Apple Silicon via LM Studio. Cuts memory footprint below the 8-bit variant at the cost of slightly higher perplexity loss. Useful when 24GB unified memory is already partially allocated by the OS or other processes.

2,054,422 ↓ · 8 ♡

Qwen3.8-27B-MLX-5bit

A 5-bit MLX quantization of Qwen3.8-27B, the most compressed publicly available MLX variant from LM Studio for this model. Fits in tighter memory envelopes at the cost of more pronounced quality degradation compared to 6-bit or 8-bit versions.

2,028,551 ↓ · 0 ♡

Qwen3.5-122B-A10B-FP8

Built for vision-language understanding, Qwen3.5-122B-A10B-FP8 is a qwen3-based model with publicly available weights. At about 122000M parameters, Qwen3.5-122B-A10B-FP8 sits in the frontier-scale tier, which sets its memory and latency budget. Qwen3.5-122B-A10B-FP8 is Apache 2.0-licensed, clearing it for closed-source and paid products. Read Qwen3.5-122B-A10B-FP8's card for hardware requirements and licensing fine print before deploying.

1,789,494 ↓ · 115 ♡

moondream2

Moondream2 is a 1.9B parameter vision-language model designed to be the smallest model that can meaningfully answer questions about images. It pairs a SigLIP vision encoder with a Phi-1.5 language backbone and achieves surprising capability at its size.

1,718,015 ↓ · 1,435 ♡

Qwen2-VL-2B-Instruct

Qwen2-VL-2B-Instruct is a 2B parameter vision-language model from Alibaba's Qwen team, supporting image and video understanding alongside text instruction-following. At 2B parameters it runs on consumer GPUs while retaining competitive OCR, chart reading, and visual QA accuracy. It is the instruction-tuned version of the Qwen2-VL-2B base.

1,652,860 ↓ · 518 ♡

Qwen2-VL-7B-Instruct-AWQ

Qwen2-VL-7B-Instruct-AWQ is a mid-sized checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps Qwen2-VL-7B-Instruct-AWQ unrestricted for commercial reuse. Weighing in near 7000M parameters, Qwen2-VL-7B-Instruct-AWQ trades some ceiling for cheaper, faster inference. Qwen2-VL-7B-Instruct-AWQ is community-maintained, so track upstream changes and pin a known-good revision.

1,615,786 ↓ · 48 ♡

Huihui-Qwen3.8-27B-abliterated-GGUF

An abliterated GGUF derivative of Qwen3.8-27B with safety layers surgically removed, packaged by huihui-ai. Abliteration edits internal refusal directions in the model's residual stream, producing a model that complies with a wider range of prompts than the original. Suitable only for isolated research environments where content controls are enforced at the application layer.

1,485,837 ↓ · 442 ♡

SmolVLM2-500M-Video-Instruct

SmolVLM2-500M-Video-Instruct is an openly licensed vision-language understanding model. At about 500M parameters, SmolVLM2-500M-Video-Instruct sits in the compact tier, which sets its memory and latency budget. SmolVLM2-500M-Video-Instruct is Apache 2.0-licensed, clearing it for closed-source and paid products. SmolVLM2-500M-Video-Instruct is community-maintained, so track upstream changes and pin a known-good revision.

1,478,731 ↓ · 173 ♡

Qwen2.5-VL-32B-Instruct

Qwen2.5-VL-32B-Instruct is Alibaba's 32B-parameter vision-language instruction model, part of the Qwen2.5-VL series described in arxiv:2502.13923. It processes image and text inputs jointly and is optimized for instruction-following tasks including document understanding, chart analysis, and visual question answering. With nearly 500 likes and broad Azure deployment support, it occupies the mid-to-large scale segment of open-weight multimodal models.

1,473,861 ↓ · 499 ♡

Qwen3.5-2B-Base

Qwen3.5-2B-Base is an uninstruct-tuned 2B parameter base model from Alibaba's Qwen team, using the Qwen3.5 architecture with image-text-to-text capability. At 2B parameters it is designed for edge deployment, fine-tuning, and low-resource inference. The Apache 2.0 license and endpoints_compatible flag make it a straightforward starting point for supervised fine-tuning pipelines.

1,469,303 ↓ · 88 ♡

diffusiongemma-26B-A4B-it

diffusiongemma-26B-A4B-it is Google's experimental diffusion-based language model built on the Gemma 4 MoE architecture, applying masked diffusion to text generation instead of autoregressive decoding. At 26B active-parameter scale it explores whether diffusion LMs can match autoregressive quality on instruction-following tasks. It accepts text and image inputs and produces text through iterative denoising.

1,447,401 ↓ · 1,194 ♡

Qwen3.5-9B-GGUF

Qwen3.5-9B-GGUF is a qwen3-based open-weight model aimed at vision-language understanding. Qwen3.5-9B-GGUF's 9000M-parameter size keeps hosting requirements modest relative to frontier models. Permissive Apache 2.0 terms let Qwen3.5-9B-GGUF go straight into commercial pipelines. Check the Qwen3.5-9B-GGUF model card for benchmarks and intended use before adopting it.

1,434,175 ↓ · 865 ♡

Qwen2.5-VL-7B-Instruct-AWQ

Qwen2.5-VL-7B-Instruct-AWQ targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. Permissive Apache 2.0 terms let Qwen2.5-VL-7B-Instruct-AWQ go straight into commercial pipelines. Prebuilt AWQ weights make local and edge inference of Qwen2.5-VL-7B-Instruct-AWQ straightforward. Like most open checkpoints, Qwen2.5-VL-7B-Instruct-AWQ rewards a quick in-domain eval before commitment.

1,396,140 ↓ · 105 ♡

gemma-3-27b-it-int4-awq

Red Hat's INT4 AWQ quantization of Google's Gemma 3 27B instruction-tuned multimodal model. AWQ (Activation-aware Weight Quantization) uses channel-wise scaling to preserve accuracy-critical weights, making it generally higher quality than naive 4-bit quantization at the same size.

1,377,294 ↓ · 40 ♡

surya-ocr-2

As an open-weight model, surya-ocr-2 focuses on vision-language understanding. surya-ocr-2 is subject to OpenRAIL terms, so confirm licensing before commercial use. Read surya-ocr-2's card for hardware requirements and licensing fine print before deploying.

1,336,357 ↓ · 99 ♡

Qwen3.6-27B-int4-AutoRound

AutoRound INT4 quantization of Qwen3.6-27B with W4G128 weight grouping and W4A16 configuration. AutoRound uses sign gradient descent to minimize quantization error, generally outperforming GPTQ at the same bit-width. Includes multi-token prediction (MTP) head for speculative decoding, which can increase throughput when paired with a draft model.

1,306,007 ↓ · 132 ♡

gemma-4-31B-it-qat-w4a16-ct

Gemma-4-31B-it-qat-w4a16-ct is a W4A16 quantization-aware trained (QAT) version of Google's Gemma 4 31B instruction-tuned multimodal model, packaged in compressed-tensors format. QAT bakes quantization into the training process rather than applying it post-hoc, generally preserving more quality than standard PTQ at the same bit width. The model handles interleaved image and text inputs for instruction-following tasks.

1,301,633 ↓ · 62 ♡

Unlimited-OCR-AWQ

An AWQ INT4 quantization of Baidu's Unlimited-OCR, a mixture-of-experts vision-language model for high-accuracy optical character recognition across multiple languages. AWQ quantization uses activation-aware scaling to reduce quality loss at INT4. The model handles complex layouts, mathematical notation, and degraded documents.

1,282,599 ↓ · 2 ♡

gemma-3-4b-it

Gemma 3 4B Instruct is Google's compact instruction-following model, targeting deployment on single-GPU and edge devices. It covers both text and image inputs and is suitable for conversational AI applications with moderate resource constraints.

1,277,061 ↓ · 1,466 ♡

Qwen3.8-27B-NVFP4

An NVFP4 (4-bit floating point) quantization of Qwen3.8-27B produced by RadixArk using NVIDIA's Model Optimizer toolkit. NVFP4 targets NVIDIA Hopper-class GPUs (H100, H200) and their TensorRT-LLM stack, offering substantial throughput gains over BF16 at the cost of some precision. Intended for production inference pipelines on NVIDIA data center hardware.

1,238,569 ↓ · 69 ♡

Qwen2-VL-7B-Instruct

Qwen2-VL 7B is Alibaba's second-generation vision-language model, instruction-tuned to follow text+image prompts. It handles variable-resolution inputs natively and scores competitively against GPT-4V on standard multimodal benchmarks at the 7B scale.

1,233,035 ↓ · 1,285 ♡

Qwen3.5-9B-AWQ

Qwen3.5-9B-AWQ is a 4-bit AWQ quantization of Qwen3.5-9B, packaged for vLLM deployment. Qwen3.5 is the multimodal variant of the Qwen3 series, and the 9B size targets a balance of quality and throughput. AWQ (Activation-aware Weight Quantization) calibrates quantization ranges to minimize output degradation, making this suitable for serving in production environments.

1,210,979 ↓ · 26 ♡

Qwen3.5-35B-A3B-FP8

Qwen3.5-35B-A3B-FP8 is an openly licensed vision-language understanding model in the qwen3 family. Prebuilt FP8 weights make local and edge inference of Qwen3.5-35B-A3B-FP8 straightforward. Qwen3.5-35B-A3B-FP8 is Apache 2.0-licensed, clearing it for closed-source and paid products. Treat Qwen3.5-35B-A3B-FP8's published metrics as a starting point and validate against your workload.

1,208,426 ↓ · 154 ♡

Qwen3.6-35B-A3B-GGUF

Unsloth's GGUF-converted and optionally quantized version of Qwen3.6-35B-A3B, optimized for local inference via llama.cpp and Ollama. Unsloth applies custom quantization recipes to reduce size while minimizing quality loss.

1,206,014 ↓ · 1,569 ♡

Qwen3.5-122B-A10B

Qwen3.5-122B-A10B targets vision-language understanding and is shipped as a frontier-scale, self-hostable checkpoint. Permissive Apache 2.0 terms let Qwen3.5-122B-A10B go straight into commercial pipelines. Qwen3.5-122B-A10B's 122000M-parameter size keeps hosting requirements modest relative to frontier models. Like most open checkpoints, Qwen3.5-122B-A10B rewards a quick in-domain eval before commitment.

1,195,228 ↓ · 611 ♡

DeepSeek-OCR-2

DeepSeek-OCR-2 is a deepseek-based open-weight model aimed at vision-language understanding. Permissive Apache 2.0 terms let DeepSeek-OCR-2 go straight into commercial pipelines. Training spans multiple languages, so DeepSeek-OCR-2 covers cross-lingual vision-language understanding from one checkpoint. Check the DeepSeek-OCR-2 model card for benchmarks and intended use before adopting it.

1,116,269 ↓ · 1,083 ♡

Qwen3.6-27B-GGUF

Qwen3.6-27B-GGUF is an openly licensed vision-language understanding model in the qwen family. Qwen3.6-27B-GGUF is Apache 2.0-licensed, clearing it for closed-source and paid products. At about 27000M parameters, Qwen3.6-27B-GGUF sits in the large tier, which sets its memory and latency budget. Evaluate Qwen3.6-27B-GGUF on your own data before trusting it in production.

1,099,191 ↓ · 947 ♡

gemma-4-31B-it-FP8-block

gemma-4-31B-it-FP8-block is a FP8 quantization for reduced VRAM on supported GPU backends (vLLM, llm-compressor) version of Google's Gemma 4 multimodal (text + image) instruction-tuned model. 31B parameters are reduced to lower-precision weights for deployment on memory-constrained hardware or Apple Silicon, with quality degradation typically small for general chat tasks. The base model is Apache-2.0 licensed.

1,069,815 ↓ · 45 ♡

Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF

HauhauCS's uncensored GGUF variant of Qwen3.8-27B with vision capability and 'Aggressive MTP' (Multi-Token Prediction) configuration. MTP speculative decoding can materially improve tokens-per-second on compatible inference servers. Safety filters removed, targeting research and creative workflows.

1,061,687 ↓ · 753 ♡

Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-GGUF

A GGUF release combining Qwen3.6-35B-A3B (MoE architecture) with uncensored fine-tuning from LuffyTheFox, incorporating the Hermes instruction style. The Genesis-Hermes-V10 naming indicates iterative instruction tuning aimed at compliance without safety filters. MoE structure keeps active parameters near 3B during inference.

1,061,057 ↓ · 564 ♡

Qwen3.5-4B-GGUF

Qwen3.5-4B-GGUF targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. Qwen3.5-4B-GGUF's 4000M-parameter size keeps hosting requirements modest relative to frontier models. Permissive Apache 2.0 terms let Qwen3.5-4B-GGUF go straight into commercial pipelines. Evaluate Qwen3.5-4B-GGUF on your own data before trusting it in production.

1,043,344 ↓ · 392 ♡

Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V9-GGUF

GGUF quantization of Qwen3.6-35B MoE abliterated and fine-tuned with Hermes function-calling data at the Genesis-Hermes-V9 merge level. Combines uncensored instruction following with structured tool-calling and vision input. Imatrix calibration applied.

1,038,728 ↓ · 540 ♡

gemma-4-26B-A4B-it-GGUF

gemma-4-26B-A4B-it-GGUF is Unsloth's GGUF quantization of Google's Gemma 4 26B mixture-of-experts instruction-tuned multimodal model. With approximately 4B active parameters per token, it runs on 16–24GB VRAM in GGUF format while retaining vision and text understanding capabilities. GGUF format provides llama.cpp and Ollama compatibility for local self-hosted deployment.

1,016,372 ↓ · 1,082 ♡

Cosmos-Reason2-2B

Cosmos-Reason2-2B is NVIDIA's 2B visual reasoning model from the Cosmos series, fine-tuned from Qwen3-VL-2B for physical world understanding tasks. It is trained to reason about spatial relationships, object interactions, and temporal dynamics in images and videos, targeting robotics and autonomous system perception research. Despite the 2B scale, the Cosmos training pipeline includes extensive world-model data.

977,013 ↓ · 169 ♡

Inkling-Small-GGUF

Unsloth GGUF quantizations of Inkling-Small, ThinkingMachines' compact multimodal MoE that accepts image, audio, and text inputs together. Q4 through Q8 variants target deployments where multimodal capability matters more than raw language benchmarks. Apache 2.0 licensed.

968,453 ↓ · 83 ♡

Muse-Glimmer-30B-GGUF

A GGUF quantization of Muse-Glimmer-30B, a 30B image-text-to-text model from the meta-models family. Unsloth provides multiple quant levels using their optimization pipeline. The model is based on architectures described in arxiv:2504.13181 and arxiv:2602.06036. Apache-2.0 licensed.

961,982 ↓ · 511 ♡

Qwen3-VL-235B-A22B-Instruct

Qwen3-VL-235B-A22B-Instruct is a frontier-scale checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps Qwen3-VL-235B-A22B-Instruct unrestricted for commercial reuse. Weighing in near 235000M parameters, Qwen3-VL-235B-A22B-Instruct trades some ceiling for cheaper, faster inference. Treat Qwen3-VL-235B-A22B-Instruct's published metrics as a starting point and validate against your workload.

951,994 ↓ · 415 ♡

gemma-3-12b-it

Gemma 3 12B is Google's mid-size instruction-tuned model in the Gemma 3 family, designed to balance capability and deployment cost. It handles text-only instruction following and is positioned between the 4B and 27B variants.

943,759 ↓ · 812 ♡

tiny-Qwen2_5_VLForConditionalGeneration

Built for vision-language understanding, tiny-Qwen2_5_VLForConditionalGeneration is a qwen2-based model with publicly available weights. Read tiny-Qwen2_5_VLForConditionalGeneration's card for hardware requirements and licensing fine print before deploying.

932,998 ↓ · 0 ♡

Qwen3.8-27B-AWQ-INT4

An AWQ (Activation-aware Weight Quantization) INT4 version of Qwen3.8-27B by cyankiwi, targeting NVIDIA GPUs via the AutoAWQ or vLLM stack. AWQ selectively quantizes weights based on activation importance, generally outperforming naive INT4 quantization in perplexity while still reducing memory footprint by ~4x versus bfloat16.

912,355 ↓ · 79 ♡

medgemma-4b-it

MedGemma-4B-it is Google's 4B instruction-tuned multimodal model specialized for medical image and text understanding, covering radiology, dermatology, pathology, and ophthalmology. It accepts medical images (chest X-rays, skin images, histology slides, fundus photos) paired with clinical questions. Not cleared for clinical decision support — research and development only.

911,590 ↓ · 1,039 ♡

Qwen3.6-27B-MTP-GGUF

Unsloth's GGUF quantisation of Qwen3.6-27B with Multi-Token Prediction (MTP) heads, enabling speculative decoding with compatible runtimes like llama.cpp. MTP allows the model to predict multiple future tokens per step, increasing throughput on CPU and single-GPU machines. Unsloth applies imatrix-based importance weighting to reduce quality loss in lower-bit GGUF variants.

891,871 ↓ · 1,306 ♡

surya-ocr-2-gguf

A GGUF conversion of Surya OCR 2, an open-source document OCR system from Datalab that handles multilingual text, column layouts, and mathematical content better than single-pass OCR models. The underlying model uses a layout-aware approach to reading order detection and text recognition. It ships under the OpenRAIL license.

812,490 ↓ · 19 ♡

Qwen3.8-27B-GGUF

The official GGUF quantization of Qwen3.8-27B produced by ggml-org, the maintainers of the llama.cpp framework. These are the reference quantizations for the model, available in multiple bit depths (Q4_K_M, Q5_K_M, Q8_0, etc.) and tested against the llama.cpp inference stack. Preferred over community quantizations when provenance and consistency matter.

796,446 ↓ · 60 ♡

Qwen3-VL-30B-A3B-Instruct-FP8

Qwen3-VL 30B MoE vision-language model in FP8 precision with 3B active parameters per token, instruction-tuned for multimodal tasks. Combines Qwen3's language capability with vision understanding, optimized for H100-class GPU serving.

795,918 ↓ · 115 ♡

Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

An imatrix GGUF quantization of Qwen3.5-9B that has been fine-tuned to remove safety alignment through the Defiant/Heretic/NEO uncensored lineage, with MTP (multi-token prediction) training applied. The 9B scale makes this runnable on consumer hardware at reasonable quant levels.

790,467 ↓ · 519 ♡

Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF

An uncensored multimodal GGUF quantization of a Qwen3.6 35B-A3B MoE derivative, incorporating the Hermes function-calling dataset to improve structured tool use and agentic behavior. The Genesis and Hermes lineage removes safety alignment from the base model; the vision capability is preserved from the Qwen3.6 series. Imatrix calibration is applied.

789,041 ↓ · 512 ♡

Qwen3.6-27B-AWQ

QuantTrio's AWQ 4-bit quantization of Qwen3.6-27B, a dense (non-MoE) multimodal model supporting image and text inputs. Tagged for vLLM serving with compressed-tensors compatibility. Qwen3.5/3.6 dense variants trade MoE routing complexity for more predictable latency.

785,755 ↓ · 24 ♡

Huihui-Qwen3.6-27B-abliterated-AWQ-MTP

4-bit AWQ quantization of a Qwen3.6-27B abliterated variant with multi-token prediction draft heads for speculative decoding. AWQ calibrates weight scale per group for better quality than naive INT4, and the MTP heads add throughput via speculation. Intended for vLLM serving.

776,301 ↓ · 13 ♡

LFM2.5-VL-450M

LFM2.5-VL-450M is LiquidAI's 450M-parameter multimodal edge model from the LFM2.5-VL family, supporting 10 languages and designed for on-device deployment on mobile and embedded hardware. It uses LiquidAI's custom LFM2 architecture (not a standard transformer) for efficient inference at the sub-500M scale. Despite the small size, it handles image+text inputs across English, Japanese, Korean, French, Spanish, German, Arabic, Chinese, Portuguese, and others.

770,595 ↓ · 194 ♡

Rax-4.5

Rax-4.5 is an image-text-to-text model from raxcore-dev built on the Qwen3.5 architecture, with conversational instruction tuning. It follows the pattern of lightweight community models fine-tuned from strong open-weight bases for specific conversational use cases. Details on training data and alignment approach are not publicly documented.

767,831 ↓ · 5 ♡

Qwen3.6-35B-A3B-AWQ

Qwen3.6-35B-A3B-AWQ targets vision-language understanding and is shipped as a frontier-scale, self-hostable checkpoint. Permissive Apache 2.0 terms let Qwen3.6-35B-A3B-AWQ go straight into commercial pipelines. Qwen3.6-35B-A3B-AWQ's 35000M-parameter size keeps hosting requirements modest relative to frontier models. Treat Qwen3.6-35B-A3B-AWQ's published metrics as a starting point and validate against your workload.

757,129 ↓ · 33 ♡

EXAONE-4.5-33B

EXAONE-4.5-33B is an open-weight vision-language understanding model. EXAONE-4.5-33B is multilingual by design rather than English-only. Licensing for EXAONE-4.5-33B is unspecified or custom — clear it before commercial use. Treat EXAONE-4.5-33B's published metrics as a starting point and validate against your workload.

757,019 ↓ · 163 ♡

UI-TARS-1.5-7B

UI-TARS-1.5-7B is ByteDance's 7B GUI agent model built on Qwen2.5-VL, fine-tuned for autonomous interaction with graphical user interfaces. It can interpret screenshots, identify UI elements, and generate action sequences (click, type, scroll) to complete computer tasks from natural language instructions. Version 1.5 improves over 1.0 on web-based task completion and cross-platform generalization.

751,000 ↓ · 596 ♡

Qwen3.6-35B-A3B-GPTQ-Int4

Qwen3.6-35B-A3B-GPTQ-Int4 is a GPTQ INT4 quantization of Alibaba's Qwen3.6-35B-A3B sparse mixture-of-experts model, which activates approximately 3B parameters per token despite a 35B total parameter count. It supports speculative decoding via MTP (multi-token prediction) and covers English, Thai, and Chinese. The quantization targets GPU memory reduction while preserving the MoE efficiency advantage of the base model.

737,881 ↓ · 30 ♡

Qwen3.6-35B-A3B-AWQ-4bit

AWQ 4-bit quantization of Qwen3.6-35B-A3B, a mixture-of-experts model that activates approximately 3B parameters per token despite 35B total parameters. The cyankiwi quantization uses compressed-tensors format compatible with vLLM. MoE architecture means memory footprint scales with total parameters, not active ones.

727,628 ↓ · 96 ♡

Qwen3.5-4B-AWQ-4bit

AWQ 4-bit quantization of Qwen3.5-4B, a dense multimodal model supporting image-text-to-text tasks. At 4B parameters with AWQ compression, inference fits within ~4 GB VRAM, making it accessible on mid-range consumer cards. compressed-tensors format targets vLLM serving.

726,371 ↓ · 21 ♡

Phi-3.5-vision-instruct

As a phi-based open-weight model, Phi-3.5-vision-instruct focuses on vision-language understanding. The MIT license keeps Phi-3.5-vision-instruct unrestricted for commercial reuse. Training spans multiple languages, so Phi-3.5-vision-instruct covers cross-lingual vision-language understanding from one checkpoint. Before relying on Phi-3.5-vision-instruct, reproduce its key numbers on representative inputs.

721,685 ↓ · 738 ♡

Qwen3.8-27B-NVFP4

Inferact's NVFP4 quantization of Qwen3.8-27B targets NVIDIA Hopper-generation GPUs for high-throughput inference. Similar to RadixArk's NVFP4 variant but produced independently. Intended for TensorRT-LLM or vLLM deployments needing 4-bit density on H100-class hardware.

719,820 ↓ · 15 ♡

Qwen3.5-397B-A17B-FP8

Qwen3.5-397B-A17B-FP8 is an openly licensed vision-language understanding model in the qwen3 family. Prebuilt FP8 weights make local and edge inference of Qwen3.5-397B-A17B-FP8 straightforward. Qwen3.5-397B-A17B-FP8 is Apache 2.0-licensed, clearing it for closed-source and paid products. Qwen3.5-397B-A17B-FP8 is community-maintained, so track upstream changes and pin a known-good revision.

709,128 ↓ · 183 ♡

gemma-4-31B

Gemma 4 31B is Google's base (non-instruct) multimodal language model with image-text-to-text capability, released under Apache 2.0. As a base model, it is intended for fine-tuning and research rather than direct deployment as a chat assistant. The 31B parameter count puts it in a tier where it competes with Mistral-Medium and Llama 3.1 70B in terms of raw capability before instruction tuning.

698,138 ↓ · 515 ♡

Kimi-K2.6

Kimi-K2.6 targets vision-language understanding and is shipped as an open-weight, self-hostable checkpoint. Licensing for Kimi-K2.6 is unspecified or custom — clear it before commercial use. Kimi-K2.6 is community-maintained, so track upstream changes and pin a known-good revision.

696,296 ↓ · 1,596 ♡

InternVL2-2B

As an internvl-based mid-sized model, InternVL2-2B focuses on vision-language understanding. Training spans multiple languages, so InternVL2-2B covers cross-lingual vision-language understanding from one checkpoint. The MIT license keeps InternVL2-2B unrestricted for commercial reuse. InternVL2-2B ships without a hosted SLA, so budget for self-managed deployment and monitoring.

692,384 ↓ · 81 ♡

GOT-OCR2_0

GOT-OCR2.0 (General OCR Theory) is a 580M-parameter image-text-to-text model from UCAS that unifies diverse OCR tasks under a single architecture. With 1,547 likes it is among the most popular specialized OCR models on HuggingFace, supporting formula, table, and scene text recognition.

688,594 ↓ · 1,557 ♡

Qwen3-VL-32B-Instruct

Qwen3-VL-32B-Instruct is a qwen3-based open-weight model aimed at vision-language understanding. Qwen3-VL-32B-Instruct's 32000M-parameter size keeps hosting requirements modest relative to frontier models. Permissive Apache 2.0 terms let Qwen3-VL-32B-Instruct go straight into commercial pipelines. Read Qwen3-VL-32B-Instruct's card for hardware requirements and licensing fine print before deploying.

687,138 ↓ · 236 ♡

gemma-4-26B-A4B-it-FP8-dynamic

Red Hat's FP8-dynamic quantization of Google's Gemma 4 26B A4B instruction-tuned multimodal model, packaged for vLLM deployment. The dynamic FP8 scheme balances throughput gains with quality retention on H100/H200 class hardware.

664,724 ↓ · 40 ♡

HunyuanOCR

As an open-weight model, HunyuanOCR focuses on vision-language understanding. HunyuanOCR lists a non-standard license, so confirm permissions before deployment. Training spans multiple languages, so HunyuanOCR covers cross-lingual vision-language understanding from one checkpoint. Check the HunyuanOCR model card for benchmarks and intended use before adopting it.

659,531 ↓ · 812 ♡

Qwen3.5-9B-NVFP4

AxionML's NVFP4 quantisation of Qwen3.5-9B using NVIDIA's ModelOpt toolkit, targeting sglang and vLLM serving on Hopper GPUs. Qwen3.5-9B is a multimodal model with image-text input capability; the NVFP4 format enables deployment at reduced memory cost while leveraging H100 4-bit tensor cores for throughput. ModelOpt-based quantisation preserves calibration-aware weight scaling.

655,146 ↓ · 18 ♡

Florence-2-large

As a florence-based open-weight model, Florence-2-large focuses on vision-language understanding. The MIT license keeps Florence-2-large unrestricted for commercial reuse. Before relying on Florence-2-large, reproduce its key numbers on representative inputs.

626,913 ↓ · 1,852 ♡

SmolVLM-256M-Instruct

SmolVLM-256M-Instruct is a compact checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps SmolVLM-256M-Instruct unrestricted for commercial reuse. Weighing in near 256M parameters, SmolVLM-256M-Instruct trades some ceiling for cheaper, faster inference. Like most open checkpoints, SmolVLM-256M-Instruct rewards a quick in-domain eval before commitment.

624,935 ↓ · 400 ♡

gemma-4-26B-A4B-it-qat-GGUF

gemma-4-26B-A4B-it-qat-GGUF is a GGUF format (quantized) for llama.cpp, LM Studio, and compatible runtimes version of Google's Gemma 4 MoE-based multimodal (text + image) instruction-tuned model. 26B parameters are reduced to lower-precision weights for deployment on memory-constrained hardware or Apple Silicon, with quality degradation typically small for general chat tasks. The base model is Apache-2.0 licensed.

618,695 ↓ · 395 ♡

gemma-4-E4B-it

Built for vision-language understanding, gemma-4-E4B-it is a gemma-based model with publicly available weights. The weights start from gemma-4-e4b-it and specialize it for the target task. gemma-4-E4B-it is Apache 2.0-licensed, clearing it for closed-source and paid products. Before relying on gemma-4-E4B-it, reproduce its key numbers on representative inputs.

618,070 ↓ · 23 ♡

InternVL2-1B

InternVL2-1B is an openly licensed vision-language understanding model in the internvl family. At about 1000M parameters, InternVL2-1B sits in the mid-sized tier, which sets its memory and latency budget. InternVL2-1B is multilingual by design rather than English-only. Like most open checkpoints, InternVL2-1B rewards a quick in-domain eval before commitment.

615,580 ↓ · 83 ♡

llava-onevision-qwen2-0.5b-ov-hf

llava-onevision-qwen2-0.5b-ov-hf is a compact checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps llava-onevision-qwen2-0.5b-ov-hf unrestricted for commercial reuse. Weighing in near 500M parameters, llava-onevision-qwen2-0.5b-ov-hf trades some ceiling for cheaper, faster inference. Treat llava-onevision-qwen2-0.5b-ov-hf's published metrics as a starting point and validate against your workload.

613,446 ↓ · 59 ♡

Qwen3.6-35B-A3B-heretic-NVFP4

NVFP4 quantization of the Heretic abliterated Qwen3.6-35B-A3B MoE for NVIDIA Blackwell GPUs. Targets production-like throughput at FP4 precision on GB10/GB200 with chunked prefill, prefix caching, and flashinfer-cutlass optimizations. Supports 1M-token context. Apache 2.0 licensed.

606,573 ↓ · 69 ♡

Kimi-K2.5

Kimi-K2.5 is an open-weight model aimed at vision-language understanding. Kimi-K2.5 lists a non-standard license, so confirm permissions before deployment. Read Kimi-K2.5's card for hardware requirements and licensing fine print before deploying.

601,096 ↓ · 2,864 ♡

Kimi-K3-GGUF

A GGUF quantization of Moonshot AI's Kimi-K3, a large multimodal conversational model. Unsloth's imatrix calibration is applied to reduce quantization error. Kimi-K3 targets complex reasoning and long-context conversational tasks.

588,843 ↓ · 372 ♡

Qwen3.5-9B-MLX-4bit

This is a 4-bit MLX-format quantization of Alibaba's Qwen3.5-9B multimodal base model, packaged by the LM Studio community for Apple Silicon deployment. MLX is Apple's machine learning framework optimized for M-series chips, so this checkpoint is specifically intended for local inference on macOS without CUDA. The underlying Qwen3.5-9B supports interleaved image and text inputs for conversational tasks.

580,277 ↓ · 4 ♡

gemma-3-27b-it-GPTQ-4b-128g

As a gemma3-based large model, gemma-3-27b-it-GPTQ-4b-128g focuses on vision-language understanding. gemma-3-27b-it-GPTQ-4b-128g is subject to Gemma terms, so confirm licensing before commercial use. GPTQ builds of gemma-3-27b-it-GPTQ-4b-128g are published alongside the full checkpoint for low-memory serving. Check the gemma-3-27b-it-GPTQ-4b-128g model card for benchmarks and intended use before adopting it.

578,678 ↓ · 45 ♡

Cosmos-Reason2-8B

As a large model, Cosmos-Reason2-8B focuses on vision-language understanding. Cosmos-Reason2-8B lists a non-standard license, so confirm permissions before deployment. The weights start from qwen3-vl-8b-instruct and specialize it for the target task. Before relying on Cosmos-Reason2-8B, reproduce its key numbers on representative inputs.

578,670 ↓ · 213 ♡

Muse-Glimmer-30B

Muse-Glimmer-30B is a 30B multimodal model from meta-models (unaffiliated with Meta AI) combining vision and text understanding. With two published arXiv papers and Azure deploy integration, it targets researchers wanting an Apache-2.0-licensed alternative to commercially-gated frontier VLMs.

570,726 ↓ · 1,806 ♡

MiniCPM-V-4.6

MiniCPM-V-4.6 is OpenBMB's MiniCPM-V 4.6, a lightweight on-device multimodal model optimized for image+text tasks at minimal parameter count. Version 4.6 targets improved document OCR, mathematical diagram understanding, and multilingual captioning within the constraints of mobile or edge deployment. It is compatible with deployment via llama.cpp or the MiniCPM-specific inference stack.

568,220 ↓ · 1,201 ♡

gemma-4-E4B-it-GGUF

Built for vision-language understanding, gemma-4-E4B-it-GGUF is a gemma-based model with publicly available weights. At about 4000M parameters, gemma-4-E4B-it-GGUF sits in the mid-sized tier, which sets its memory and latency budget. GGUF builds of gemma-4-E4B-it-GGUF are published alongside the full checkpoint for low-memory serving. Before relying on gemma-4-E4B-it-GGUF, reproduce its key numbers on representative inputs.

564,258 ↓ · 598 ♡

Qwen3.6-27B-MLX-8bit

Qwen3.6-27B-MLX-8bit is an MLX 8-bit quantized version of Qwen3.6-27B, packaged by lmstudio-community for inference on Apple Silicon via the MLX framework. MLX quantization converts the model to integer weights while preserving floating-point activations, enabling the 27B model to run within 16-24GB unified memory on M2/M3 Pro or Ultra configurations. Intended for use in LM Studio or direct MLX inference.

563,830 ↓ · 16 ♡

endless-frontier_BigBang-v1-GGUF

BigBang-v1 is an image-text-to-text model from XYZAILab's endless-frontier series, quantized to GGUF by bartowski. Bartowski is a well-known community quantizer providing consistent imatrix-calibrated GGUF variants across many models. This is a vision-language model supporting multimodal chat via standard llama.cpp tooling.

561,794 ↓ · 28 ♡

gemma-4-31B-it-qat-GGUF

Unsloth's GGUF-format conversion of Google's Gemma-4-31B QAT (Quantization-Aware Training) checkpoint. QAT weights are pre-quantized during training rather than post-hoc, which typically preserves more accuracy than PTQ at the same bit width. Compatible with llama.cpp for CPU and hybrid inference.

561,084 ↓ · 184 ♡

gemma-4-E2B-it-GGUF

Built for vision-language understanding, gemma-4-E2B-it-GGUF is a gemma-based model with publicly available weights. gemma-4-E2B-it-GGUF is Apache 2.0-licensed, clearing it for closed-source and paid products. At about 2000M parameters, gemma-4-E2B-it-GGUF sits in the mid-sized tier, which sets its memory and latency budget. gemma-4-E2B-it-GGUF ships without a hosted SLA, so budget for self-managed deployment and monitoring.

553,797 ↓ · 294 ♡

Qwen3.6-27B-MLX-4bit

Qwen3.6-27B-MLX-4bit is an MLX 4-bit quantized version of Qwen3.6-27B, packaged by lmstudio-community for inference on Apple Silicon via the MLX framework. MLX quantization converts the model to integer weights while preserving floating-point activations, enabling the 27B model to run within 16-24GB unified memory on M2/M3 Pro or Ultra configurations. Intended for use in LM Studio or direct MLX inference.

550,071 ↓ · 16 ♡

Qwen3.5-9B-MLX-8bit

Qwen3.5-9B-MLX-8bit is an 8-bit quantized version of Alibaba's Qwen3.5-9B multimodal model, packaged for Apple Silicon via the MLX framework. It handles interleaved image and text inputs for conversational use cases, making it accessible on consumer Mac hardware without requiring a discrete GPU.

548,877 ↓ · 1 ♡

Qwen3-VL-4B-Thinking

Qwen3-VL-4B-Thinking targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. Permissive Apache 2.0 terms let Qwen3-VL-4B-Thinking go straight into commercial pipelines. Qwen3-VL-4B-Thinking's 4000M-parameter size keeps hosting requirements modest relative to frontier models. Qwen3-VL-4B-Thinking is community-maintained, so track upstream changes and pin a known-good revision.

542,482 ↓ · 111 ♡

Qwen3.5-0.8B-Base

Qwen3.5-0.8B-Base is Alibaba's sub-billion-parameter base language model from the Qwen3.5 family. At 0.8B parameters, it is optimized for edge deployment and lightweight inference while retaining the architectural improvements from the Qwen3.5 series.

528,739 ↓ · 97 ♡

Florence-2-large

Florence-2-large is Microsoft's 0.77B vision foundation model that unifies multiple vision tasks — captioning, object detection, phrase grounding, and OCR — into a single sequence-to-sequence architecture with task prompts. It is described in arxiv:2311.06242 and trained on the FLD-5B dataset of 5 billion annotations.

525,864 ↓ · 6 ♡

Qwen3.6-27B-MLX-6bit

Qwen3.6-27B-MLX-6bit is an MLX 6-bit quantized version of Qwen3.6-27B, packaged by lmstudio-community for inference on Apple Silicon via the MLX framework. MLX quantization converts the model to integer weights while preserving floating-point activations, enabling the 27B model to run within 16-24GB unified memory on M2/M3 Pro or Ultra configurations. Intended for use in LM Studio or direct MLX inference.

521,063 ↓ · 2 ♡

Qwen3.6-27B-MLX-5bit

Qwen3.6-27B-MLX-5bit is an MLX 5-bit quantized version of Qwen3.6-27B, packaged by lmstudio-community for inference on Apple Silicon via the MLX framework. MLX quantization converts the model to integer weights while preserving floating-point activations, enabling the 27B model to run within 16-24GB unified memory on M2/M3 Pro or Ultra configurations. Intended for use in LM Studio or direct MLX inference.

519,767 ↓ · 0 ♡

Qwen3-VL-32B-Instruct-FP8

Qwen3-VL-32B-Instruct-FP8 targets vision-language understanding and is shipped as a frontier-scale, self-hostable checkpoint. Qwen3-VL-32B-Instruct-FP8's 32000M-parameter size keeps hosting requirements modest relative to frontier models. Permissive Apache 2.0 terms let Qwen3-VL-32B-Instruct-FP8 go straight into commercial pipelines. Treat Qwen3-VL-32B-Instruct-FP8's published metrics as a starting point and validate against your workload.

504,130 ↓ · 49 ♡

gemma-3-27b-it-FP8-dynamic

This is a dynamically quantized FP8 variant of Google's Gemma-3-27B-IT, packaged by Red Hat AI for deployment with vLLM and compressed-tensors. The quantization reduces memory footprint compared to the BF16 base, enabling the 27B model to run on hardware that would otherwise require multi-GPU setups. It is derived from the base Gemma-3-27B-IT weights and inherits both its vision capabilities and Apache 2.0 license.

500,227 ↓ · 13 ♡

gemma-4-E4B-it-unsloth-bnb-4bit

gemma-4-E4B-it-unsloth-bnb-4bit is a 4-bit bitsandbytes quantization of Google's Gemma-4-E4B instruction-tuned model, packaged by the Unsloth team. Unsloth quantizations are commonly used to enable fine-tuning and inference of Gemma models on consumer GPUs with limited VRAM. It inherits the Apache 2.0 license from the base model.

498,615 ↓ · 22 ♡

Mage-VL

Mage-VL is Microsoft's multimodal vision-language model designed for streaming video understanding alongside image and text inputs. Published alongside arxiv:2607.24904, it uses a custom mage_vl architecture with conversational capabilities. The Apache 2.0 license and 440K downloads suggest active adoption for video-QA research tasks.

497,690 ↓ · 389 ♡

gemma-4-31B-it-GGUF

gemma-4-31B-it-GGUF targets vision-language understanding and is shipped as a large, self-hostable checkpoint. gemma-4-31B-it-GGUF's 31000M-parameter size keeps hosting requirements modest relative to frontier models. Prebuilt GGUF weights make local and edge inference of gemma-4-31B-it-GGUF straightforward. gemma-4-31B-it-GGUF is community-maintained, so track upstream changes and pin a known-good revision.

493,264 ↓ · 594 ♡

gemma-4-26B-A4B-it-QAT-MLX-4bit

gemma-4-26B-A4B-it-QAT-MLX-4bit is a MLX 4-bit quantized weights optimized for Apple Silicon inference version of Google's Gemma 4 MoE-based multimodal (text + image) instruction-tuned model. 26B parameters are reduced to lower-precision weights for deployment on memory-constrained hardware or Apple Silicon, with quality degradation typically small for general chat tasks. The base model is Apache-2.0 licensed.

492,806 ↓ · 15 ♡

Qwen3.6-35B-A3B-MTP-GGUF

Unsloth's GGUF quantisation of Qwen3.6-35B, a sparse MoE model with 35B total parameters but only ~3B active per token, enhanced with Multi-Token Prediction heads for speculative decoding. The imatrix calibration in Unsloth's quantisation pipeline reduces perplexity loss compared to uncalibrated GGUF. At 35B total capacity this is a large multimodal model that fits on consumer hardware only through aggressive quantisation.

481,522 ↓ · 883 ♡

Muse-Glimmer-30B-GGUF

GGUF quantization of Muse-Glimmer-30B, a 30B multimodal model from meta-models (unaffiliated with Meta AI). Covers standard quantization levels for llama.cpp compatibility. Apache 2.0 licensed, with two arXiv papers cited in the model card.

481,393 ↓ · 322 ♡

gemma-4-12b-it-GGUF

gemma-4-12b-it-GGUF is a GGUF format (quantized) for llama.cpp, LM Studio, and compatible runtimes version of Google's Gemma 4 multimodal (text + image) instruction-tuned model. parameters are reduced to lower-precision weights for deployment on memory-constrained hardware or Apple Silicon, with quality degradation typically small for general chat tasks. The base model is Apache-2.0 licensed.

481,073 ↓ · 811 ♡

Qwen3.5-122B-A10B-GPTQ-Int4

Built for vision-language understanding, Qwen3.5-122B-A10B-GPTQ-Int4 is a qwen3-based model with publicly available weights. GPTQ builds of Qwen3.5-122B-A10B-GPTQ-Int4 are published alongside the full checkpoint for low-memory serving. Qwen3.5-122B-A10B-GPTQ-Int4 is Apache 2.0-licensed, clearing it for closed-source and paid products. Check the Qwen3.5-122B-A10B-GPTQ-Int4 model card for benchmarks and intended use before adopting it.

479,806 ↓ · 51 ♡

Molmo2-8B

Built for vision-language understanding, Molmo2-8B is an olmo-based model with publicly available weights. At about 8000M parameters, Molmo2-8B sits in the large tier, which sets its memory and latency budget. Molmo2-8B is Apache 2.0-licensed, clearing it for closed-source and paid products. Before relying on Molmo2-8B, reproduce its key numbers on representative inputs.

478,310 ↓ · 189 ♡

XYZAILab_XYZ-Aquila-mini-GGUF

XYZ-Aquila-mini from XYZAILab, quantized to GGUF by bartowski, is a compact agentic search-oriented image-text-to-text model. The Qwen3.6 base and agentic-search tags suggest it was fine-tuned for tool-calling and web search integration tasks alongside vision capability. Bartowski's imatrix calibration produces higher-quality quantizations than standard K-quant alone.

477,933 ↓ · 3 ♡

Qwen3.5-122B-A10B-GGUF

Unsloth's GGUF quantization of Qwen3.5-122B-A10B, a very large MoE model with 122B total parameters and ~10B active per token. Unsloth specializes in memory-efficient model packaging with dynamic quantization support. At 122B total parameters, this model requires substantial RAM (40–80GB) even in 4-bit GGUF, making it suited for high-end workstations or small server clusters.

472,824 ↓ · 297 ♡

Qwythos-9B-v2-GGUF

GGUF quantization of Qwythos-9B-v2, a multimodal image-text model from empero-ai. At 9B parameters it targets consumer and hobbyist users running image-plus-text tasks locally via llama.cpp-compatible runtimes. Apache 2.0 licensed.

472,364 ↓ · 261 ♡

gemma-4-26B-A4B-it-qat-q4_0-gguf

gemma-4-26B-A4B-it-qat-q4_0-gguf targets vision-language understanding and is shipped as a large, self-hostable checkpoint. Prebuilt GGUF weights make local and edge inference of gemma-4-26B-A4B-it-qat-q4_0-gguf straightforward. gemma-4-26B-A4B-it-qat-q4_0-gguf's 26000M-parameter size keeps hosting requirements modest relative to frontier models. Like most open checkpoints, gemma-4-26B-A4B-it-qat-q4_0-gguf rewards a quick in-domain eval before commitment.

470,032 ↓ · 156 ♡

Qwen3.6-35B-A3B-Uncensored-Wasserstein-GGUF

An abliterated (refusal-removed) GGUF fine-tune of Qwen3.6-35B-A3B, produced via a Wasserstein-distance-guided weight adjustment technique to remove model refusal behaviour. The 'uncensored' label means safety filters have been deliberately removed. This is a community research release targeting users who need the model to engage with content that safety-tuned models decline.

468,102 ↓ · 96 ♡

inkling-GGUF

A GGUF quantization of Inkling by Thinking Machines, a multimodal MoE language model. Unsloth's conversion provides multiple quantization levels with imatrix calibration. Inkling targets instruction-following and reasoning tasks in a mixture-of-experts architecture.

465,779 ↓ · 134 ♡

llava-v1.6-mistral-7b-hf

llava-v1.6-mistral-7b-hf targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. llava-v1.6-mistral-7b-hf's 7000M-parameter size keeps hosting requirements modest relative to frontier models. Permissive Apache 2.0 terms let llava-v1.6-mistral-7b-hf go straight into commercial pipelines. llava-v1.6-mistral-7b-hf is community-maintained, so track upstream changes and pin a known-good revision.

462,465 ↓ · 314 ♡

gemma-3-27b-it

gemma-3-27b-it is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. It is a fine-tune of gemma-3-27b-pt, inheriting that base model's general competence. Weighing in near 27000M parameters, gemma-3-27b-it trades some ceiling for cheaper, faster inference. Like most open checkpoints, gemma-3-27b-it rewards a quick in-domain eval before commitment.

448,939 ↓ · 2,014 ♡

AIN

AIN is an open-weight checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The MIT license keeps AIN unrestricted for commercial reuse. Evaluate AIN on your own data before trusting it in production.

447,167 ↓ · 21 ♡

Qwen3-VL-4B-Instruct-FP8

Qwen3-VL-4B-Instruct-FP8 is an FP8-quantized vision-language model from Alibaba, derived from the Qwen3-VL-4B-Instruct base and designed for image-grounded conversation. At 4B active parameters it targets deployments where GPU memory is constrained but multimodal instruction following is required. The model is backed by multiple arXiv publications covering the VL and Qwen3 training methodologies.

447,030 ↓ · 63 ♡

LightOnOCR-2-1B

LightOnOCR-2-1B is an open-weight model aimed at vision-language understanding. Training spans multiple languages, so LightOnOCR-2-1B covers cross-lingual vision-language understanding from one checkpoint. LightOnOCR-2-1B's 1000M-parameter size keeps hosting requirements modest relative to frontier models. LightOnOCR-2-1B ships without a hosted SLA, so budget for self-managed deployment and monitoring.

442,132 ↓ · 798 ♡

InternVL2_5-4B

InternVL2.5-4B is a 4B vision-language model combining InternViT-300M with a Qwen2.5-3B language backbone, designed for competitive multimodal benchmarks at a small parameter count. The model covers image understanding, video, and document parsing tasks. Multiple papers document the InternVL family: arxiv:2312.14238, 2404.16821, 2410.16261, 2412.05271.

441,887 ↓ · 58 ♡

gemma-3-4b-pt

Gemma-3-4B-PT is Google's 4B pre-trained (base, non-instruction-tuned) multimodal model from the Gemma-3 family, accepting image and text inputs. As a base model it requires fine-tuning or careful prompting for task-specific use. The gemma license applies.

440,651 ↓ · 155 ♡

Kimi-K2.7-Code

Kimi-K2.7-Code is Moonshot AI's code-focused multimodal model built on the kimi_k25 architecture, accepting both image and text inputs. It uses compressed-tensors for efficient weight storage and exposes custom model code, indicating non-standard architectural components beyond base Transformers.

440,001 ↓ · 1,370 ♡

medgemma-27b-it

medgemma-27b-it is Google's 27B medical vision-language model, fine-tuned from Gemma 3 on radiology reports, chest X-rays, histopathology slides, dermatology images, and ophthalmology fundus photographs. It is designed for medical image interpretation research, not clinical deployment. The model accepts image+text input and outputs clinical-style text descriptions, differentials, or structured findings.

439,019 ↓ · 382 ♡

Qwen3-VL-2B-Instruct-GGUF

Qwen3-VL-2B-Instruct-GGUF is Unsloth's GGUF distribution of Qwen3-VL-2B-Instruct, making the 2B vision-language model directly usable in llama.cpp, LM Studio, and Ollama. At 2B parameters, it is intended for on-device or memory-constrained deployment scenarios where a capable VLM must run locally. Quantization options (Q4, Q5, Q8) allow further trade-offs between quality and memory.

438,837 ↓ · 38 ♡

SmolVLM-500M-Instruct

SmolVLM-500M-Instruct is an openly licensed vision-language understanding model. At about 500M parameters, SmolVLM-500M-Instruct sits in the compact tier, which sets its memory and latency budget. SmolVLM-500M-Instruct is Apache 2.0-licensed, clearing it for closed-source and paid products. SmolVLM-500M-Instruct is community-maintained, so track upstream changes and pin a known-good revision.

435,660 ↓ · 196 ♡

Qwen3.5-35B-A3B-GPTQ-Int4

Qwen3.5-35B-A3B-GPTQ-Int4 is a frontier-scale checkpoint for vision-language understanding, distributed on the HuggingFace Hub. Weighing in near 35000M parameters, Qwen3.5-35B-A3B-GPTQ-Int4 trades some ceiling for cheaper, faster inference. The Apache 2.0 license keeps Qwen3.5-35B-A3B-GPTQ-Int4 unrestricted for commercial reuse. Qwen3.5-35B-A3B-GPTQ-Int4 is community-maintained, so track upstream changes and pin a known-good revision.

435,453 ↓ · 91 ♡

deepseek-vl2-tiny

deepseek-vl2-tiny is a deepseek-based open-weight model aimed at vision-language understanding. deepseek-vl2-tiny lists a non-standard license, so confirm permissions before deployment. deepseek-vl2-tiny ships without a hosted SLA, so budget for self-managed deployment and monitoring.

435,131 ↓ · 249 ♡

Qwen_Qwen3.6-35B-A3B-GGUF

Bartowski's GGUF quantization of Qwen3.6-35B-A3B, a mixture-of-experts multimodal model with 35B total and roughly 3.6B active parameters per forward pass. Supports image-text-to-text tasks and runs via llama.cpp on consumer hardware. Apache 2.0 licensed.

434,104 ↓ · 146 ♡

medgemma-1.5-4b-it

MedGemma-1.5-4B-IT is Google's 4-billion-parameter multimodal model fine-tuned for medical imaging and clinical reasoning tasks, built on the Gemma 3 architecture. It targets radiology, dermatology, pathology, and ophthalmology use cases, including chest X-ray interpretation and conversational clinical support. The model ships under a non-standard license that restricts certain commercial deployments in medical contexts.

433,282 ↓ · 759 ♡

blip2-opt-2.7b

blip2-opt-2.7b is a blip-based open-weight model aimed at vision-language understanding. Permissive MIT terms let blip2-opt-2.7b go straight into commercial pipelines. blip2-opt-2.7b's 2700M-parameter size keeps hosting requirements modest relative to frontier models. Check the blip2-opt-2.7b model card for benchmarks and intended use before adopting it.

431,674 ↓ · 447 ♡

Ornith-1.0-35B-NVFP4

DeepReinforce AI's NVFP4-quantized 35B MoE model from the Ornith 1.0 series, targeting agentic coding and coding-assistant use cases. The agentic-coding tag indicates training specifically for tool-use and multi-step code generation.

429,935 ↓ · 25 ♡

InternVL2-8B

InternVL2-8B is an internvl-based open-weight model aimed at vision-language understanding. Permissive MIT terms let InternVL2-8B go straight into commercial pipelines. Training spans multiple languages, so InternVL2-8B covers cross-lingual vision-language understanding from one checkpoint. Before relying on InternVL2-8B, reproduce its key numbers on representative inputs.

429,442 ↓ · 187 ♡

Qwen3-VL-8B-Instruct-AWQ-4bit

This is a 4-bit AWQ quantization of Qwen/Qwen3-VL-8B-Instruct, packaged with compressed-tensors by community contributor cyankiwi. The Qwen3-VL architecture supports vision-language instruction following, and the AWQ quantization brings the model into consumer GPU range. Based on arxiv papers 2505.09388, 2502.13923, and others documenting the Qwen VL lineage.

428,181 ↓ · 17 ♡

Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive

Imatrix-calibrated GGUF quantizations of an aggressively abliterated Qwen3.5-122B-A10B MoE with vision capability. HauhauCS's Aggressive series applies multiple rounds of refusal removal. Multilingual English and Chinese. Apache 2.0 base model license.

427,132 ↓ · 173 ♡

Qwen3.5-27B-FP8

Qwen3.5-27B-FP8 targets vision-language understanding and is shipped as a large, self-hostable checkpoint. Qwen3.5-27B-FP8's 27000M-parameter size keeps hosting requirements modest relative to frontier models. Prebuilt FP8 weights make local and edge inference of Qwen3.5-27B-FP8 straightforward. Qwen3.5-27B-FP8 is community-maintained, so track upstream changes and pin a known-good revision.

424,513 ↓ · 137 ♡

Qwen3.6-35B-A3B-NVFP4-Fast

This Unsloth repack of Qwen3.6-35B-A3B applies NVFP4 quantization via compressed-tensors for deployment on NVIDIA's latest GPUs. The '35B-A3B' label indicates a 35B total / 3B active parameter MoE configuration, making it extremely efficient per inference call. The NVFP4-Fast variant is designed for maximum throughput rather than minimum latency.

421,511 ↓ · 111 ♡

Qwen3.5-2B-AWQ-4bit

An AWQ 4-bit quantized version of Qwen/Qwen3.5-2B, adapted for image-text-to-text conversational use. The compressed-tensors format targets inference frameworks that support quantized loading such as vLLM. At 2B parameters in 4-bit precision, this checkpoint is aimed at deployments where memory is constrained.

418,098 ↓ · 3 ♡

Qwen2.5-VL-7B-Instruct-GGUF

Qwen2.5-VL-7B-Instruct-GGUF targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. Permissive Apache 2.0 terms let Qwen2.5-VL-7B-Instruct-GGUF go straight into commercial pipelines. Qwen2.5-VL-7B-Instruct-GGUF's 7000M-parameter size keeps hosting requirements modest relative to frontier models. Treat Qwen2.5-VL-7B-Instruct-GGUF's published metrics as a starting point and validate against your workload.

414,170 ↓ · 197 ♡

Nanonets-OCR2-3B

As a mid-sized model, Nanonets-OCR2-3B focuses on vision-language understanding. Training spans multiple languages, so Nanonets-OCR2-3B covers cross-lingual vision-language understanding from one checkpoint. The weights start from qwen2.5-vl-3b-instruct and specialize it for the target task. Nanonets-OCR2-3B ships without a hosted SLA, so budget for self-managed deployment and monitoring.

413,995 ↓ · 514 ♡

MiniMax-M3-MXFP8

MiniMax-M3-MXFP8 is an MXFP8-quantized variant of MiniMaxAI's MiniMax-M3 multimodal mixture-of-experts model, described in arxiv:2606.13392. It supports image, text, and video inputs and is designed for agent and coding workflows. The MXFP8 quantization reduces memory footprint compared to the full-precision base while targeting hardware that supports the MX floating-point standard.

413,032 ↓ · 54 ♡

dots.mocr

dots.mocr is an OCR-focused image-text model from dots-studio, using a custom dots_mocr architecture alongside a dot_ocr text-generation backend. It targets document and scene text recognition tasks, outputting extracted text from images. The dots_mocr architecture is specific to dots-studio and not documented in external literature.

412,178 ↓ · 163 ♡

Llama-4-Scout-17B-16E-Instruct

Llama 4 Scout is Meta's first MoE entry in the Llama series: 17B parameters per expert across 16 experts, with a small number active per token. The instruct variant follows instructions and handles image-text inputs natively, supporting 12 languages. Scout targets deployments where multimodal capability is needed at a lower active-parameter cost than dense Llama 3 models.

411,931 ↓ · 1,333 ♡

Qwen3.6-35B-A3B-MLX-4bit

Qwen3.6-35B-A3B-MLX-4bit is a 4-bit MLX-format quantization of Qwen3.6-35B-A3B, a mixture-of-experts image-text-to-text model targeting Apple Silicon inference via the MLX framework. With 35B total parameters but only ~3B active per forward pass, it is designed for efficient multimodal inference on M-series Macs. The model is Apache-2.0 licensed.

411,834 ↓ · 10 ♡

Qwen3.5-35B-A3B-AWQ-4bit

As a qwen3-based frontier-scale model, Qwen3.5-35B-A3B-AWQ-4bit focuses on vision-language understanding. Weighing in near 35000M parameters, Qwen3.5-35B-A3B-AWQ-4bit trades some ceiling for cheaper, faster inference. The Apache 2.0 license keeps Qwen3.5-35B-A3B-AWQ-4bit unrestricted for commercial reuse. Qwen3.5-35B-A3B-AWQ-4bit ships without a hosted SLA, so budget for self-managed deployment and monitoring.

411,654 ↓ · 46 ♡

Infinity-Parser2-Pro

Infinity-Parser2-Pro converts complex documents to structured Markdown, handling mixed-content pages including tables, charts, LaTeX formulas, and chemical structures via a Qwen3.5 MoE backbone. English and Chinese supported with published benchmarks on arXiv:2607.07836.

410,874 ↓ · 90 ♡

Qwen3-VL-30B-A3B-Instruct

Qwen3-VL-30B-A3B-Instruct is an openly licensed vision-language understanding model in the qwen3 family. At about 30000M parameters, Qwen3-VL-30B-A3B-Instruct sits in the large tier, which sets its memory and latency budget. Qwen3-VL-30B-A3B-Instruct is Apache 2.0-licensed, clearing it for closed-source and paid products. Like most open checkpoints, Qwen3-VL-30B-A3B-Instruct rewards a quick in-domain eval before commitment.

409,897 ↓ · 595 ♡

Qwopus3.6-35B-A3B-Coder-MTP-GGUF

Bartowski's GGUF packaging of Qwopus3.6-35B, a coding-focused merge combining Qwen3.6-35B's MoE architecture with MTP (multi-token prediction) capability. Targets llama.cpp users who need a 35B-scale coding model for local inference.

409,608 ↓ · 235 ♡

MinerU2.5-2509-1.2B

Built for vision-language understanding, MinerU2.5-2509-1.2B is a model with publicly available weights. Distribution of MinerU2.5-2509-1.2B is under AGPL-3.0, which is worth reading before you ship. At about 1200M parameters, MinerU2.5-2509-1.2B sits in the mid-sized tier, which sets its memory and latency budget. Read MinerU2.5-2509-1.2B's card for hardware requirements and licensing fine print before deploying.

409,174 ↓ · 356 ♡

Qwen3.5-9B-Base

Alibaba's 9B-parameter base language model from the Qwen3.5 series, before instruction tuning. The base model is trained on general text data and serves as a starting point for supervised fine-tuning or RLHF alignment pipelines.

408,967 ↓ · 102 ♡

granite-docling-258M

granite-docling-258M is a 258M-parameter vision-language model fine-tuned specifically for document understanding tasks within the Docling pipeline. It handles OCR, layout parsing, table extraction, formula recognition, and chart reading in a single inference pass. The model is built on the Idefics3 architecture and integrates directly with the open-source Docling library.

407,501 ↓ · 1,254 ♡

Llama-3.1-Nemotron-Nano-VL-8B-V1

Llama-3.1-Nemotron-Nano-VL-8B-V1 is an open-weight vision-language understanding model in the nemotron family. At about 8000M parameters, Llama-3.1-Nemotron-Nano-VL-8B-V1 sits in the large tier, which sets its memory and latency budget. Licensing for Llama-3.1-Nemotron-Nano-VL-8B-V1 is unspecified or custom — clear it before commercial use. Like most open checkpoints, Llama-3.1-Nemotron-Nano-VL-8B-V1 rewards a quick in-domain eval before commitment.

403,980 ↓ · 181 ♡

Huihui-Qwen3.6-35B-A3B-abliterated-FP8-DYNAMIC

FP8-dynamic quantization of huihui-ai's abliterated variant of Qwen3.6-35B-A3B. Abliteration removes safety refusal behaviors by suppressing the refusal direction in the model's residual stream during inference. Dynamic FP8 quantization allows deployment without a calibration dataset. Apache 2.0 licensed.

403,529 ↓ · 4 ♡

Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP

HauhauCS's uncensored fine-tune of Gemma 4's 26B MoE (4B active) variant with QAT (Quantization-Aware Training) and MTP (Multi-Token Prediction) configuration, in GGUF format. Combines Google's QAT quantization quality with HauhauCS's safety-removal fine-tuning and speculative decoding optimization. The 'Balanced' designation contrasts with more aggressive MTP settings.

402,068 ↓ · 198 ♡

Qwen3.5-397B-A17B

As a qwen3-based frontier-scale model, Qwen3.5-397B-A17B focuses on vision-language understanding. Weighing in near 397000M parameters, Qwen3.5-397B-A17B trades some ceiling for cheaper, faster inference. The Apache 2.0 license keeps Qwen3.5-397B-A17B unrestricted for commercial reuse. Before relying on Qwen3.5-397B-A17B, reproduce its key numbers on representative inputs.

401,650 ↓ · 1,548 ♡

GLM-4.6V-Flash

GLM-4.6V-Flash is ZhipuAI's Flash-tier vision-language model supporting Chinese and English input, designed for low-latency multimodal inference. As a Flash variant in the GLM-4 family it trades some accuracy for reduced response latency relative to the full GLM-4V.

401,356 ↓ · 619 ♡

gemma-4-26B-A4B

gemma-4-26B-A4B is a mixture-of-experts image-text-to-text model from Google with 26B total parameters and approximately 4B active parameters per forward pass. It is part of the Gemma 4 model family and supports multimodal inputs. The model is released under Apache 2.0 and is compatible with Transformers and HuggingFace inference endpoints.

400,805 ↓ · 377 ♡

Qwen3-VL-2B-Instruct-AWQ-4bit

Qwen3-VL-2B-Instruct-AWQ-4bit is a mid-sized checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps Qwen3-VL-2B-Instruct-AWQ-4bit unrestricted for commercial reuse. Prebuilt AWQ/4BIT weights make local and edge inference of Qwen3-VL-2B-Instruct-AWQ-4bit straightforward. Treat Qwen3-VL-2B-Instruct-AWQ-4bit's published metrics as a starting point and validate against your workload.

400,633 ↓ · 1 ♡

Qwen3-VL-32B-Instruct-bnb-4bit

Unsloth's bitsandbytes 4-bit quantization of Qwen3-VL-32B-Instruct, optimized for memory-efficient fine-tuning and inference. Unsloth's bnb integration targets users who want to run or fine-tune a 32B vision-language model on consumer-grade hardware.

397,806 ↓ · 5 ♡

ThinkingCap-Qwen3.6-27B-GGUF

Bartowski's GGUF conversion of ThinkingCap-Qwen3.6-27B, a model fine-tuned on efficient thinking traces that reduce unnecessary chain-of-thought tokens. It targets llama.cpp users who want reasoning capability without the token overhead of full QwQ-style outputs.

395,303 ↓ · 230 ♡

Kimi-K2.7-Code-GGUF

Kimi-K2.7-Code-GGUF is Unsloth's quantized GGUF export of Moonshot AI's Kimi-K2.7-Code, a model built for software development tasks including code completion and generation. The kimi_k25 architecture is proprietary with limited external documentation.

394,918 ↓ · 182 ♡

dots.ocr

dots.ocr is RedNote's specialized OCR model for structured document parsing, capable of extracting text from complex layouts including tables, mathematical formulas, and mixed Chinese-English documents. The 1315 community likes reflect substantial real-world adoption for document digitization use cases.

394,815 ↓ · 1,315 ♡

Qwopus3.6-27B-Coder-Compat-MTP-GGUF

Qwopus3.6-27B-Coder-Compat-MTP-GGUF is a community fine-tune of Qwen3.6-27B optimized for code generation, enhanced with a Multi-Token Prediction head for speculative decoding, and packaged as GGUF. It supports multimodal inputs and targets multilingual code tasks across five languages.

394,323 ↓ · 132 ♡

Kimi-VL-A3B-Instruct

Moonshot AI's Kimi-VL-A3B-Instruct is a 3B-active-parameter (from a larger MoE pool) vision-language model with long-context support, video understanding, and screen-grounding capabilities. It is optimized for agentic tasks like GUI navigation and spatial reasoning.

394,242 ↓ · 280 ♡

Qwen3.6-27B-AWQ-BF16-INT4

Qwen3.6-27B-AWQ-BF16-INT4 is an AWQ INT4 quantization of Alibaba's Qwen3.6-27B multimodal model, supporting both image and text inputs. The compressed-tensors format is used for weight storage, and the model targets deployments where the full BF16 27B weight set exceeds available VRAM. Being a community quantization, it inherits the Apache 2.0 license from the base model.

393,273 ↓ · 42 ♡

Qwen3-VL-32B-Instruct-AWQ

Alibaba's AWQ-quantized Qwen3-VL-32B-Instruct, a 32B vision-language model supporting image and text input. AWQ's activation-aware channel scaling minimizes accuracy loss at 4-bit weight compression for this large multimodal model.

392,894 ↓ · 12 ♡

gemma-4-31B-it-qat-q4_0-gguf

gemma-4-31B-it-qat-q4_0-gguf is an open-source image-text-to-text model available on HuggingFace. Details are sourced from the public model registry.

392,890 ↓ · 127 ♡

MinerU2.5-Pro-2604-1.2B

MinerU2.5-Pro is a 1.2B-parameter vision-language model from OpenDataLab fine-tuned for high-accuracy document parsing, built on the Qwen2-VL backbone. It handles mixed Chinese-English documents, extracting structured content from PDFs including formulas, tables, and figures. Version 2.5-Pro targets production-grade accuracy improvements over the earlier MinerU pipeline.

392,153 ↓ · 158 ♡

Qwen3.5-27B-GGUF

Qwen3.5-27B-GGUF is a qwen3-based open-weight model aimed at vision-language understanding. Qwen3.5-27B-GGUF's 27000M-parameter size keeps hosting requirements modest relative to frontier models. GGUF builds of Qwen3.5-27B-GGUF are published alongside the full checkpoint for low-memory serving. Check the Qwen3.5-27B-GGUF model card for benchmarks and intended use before adopting it.

390,551 ↓ · 503 ♡

Qwen2.5-VL-72B-Instruct

Qwen2.5-VL-72B-Instruct is Qwen's 72B vision-language model, the largest in the Qwen2.5-VL series, handling image, video, and text inputs with a 32K token context window. At 72B scale it targets document understanding, complex visual reasoning, and structured extraction from multi-page documents. It supports bounding box output for grounded visual answers.

389,793 ↓ · 649 ♡

Qwen3.5-27B-AWQ-4bit

A 4-bit AWQ quantisation of Qwen3.5-27B, a multimodal model combining image and text understanding at 27B parameters. AWQ preserves the most activationally important weights at higher precision, minimising accuracy loss compared to round-to-nearest quantisation. The result fits in significantly less GPU memory than the BF16 checkpoint while remaining compatible with vLLM and transformers backends.

389,522 ↓ · 42 ♡

Qwen3.6-27B-Uncensored-HauhauCS-Balanced

HauhauCS's balanced uncensored variant of Qwen3.6-27B, distributed in GGUF format with vision and multimodal support. The 'Balanced' designation means less aggressive safety removal compared to HauhauCS's 'Aggressive' variants, attempting to preserve some quality that full abliteration can damage while still enabling compliance with a wider range of prompts.

389,076 ↓ · 206 ♡

Qwen3.5-4B-Base

Qwen3.5-4B-Base is an open-source image-text-to-text model available on HuggingFace. Details are sourced from the public model registry.

388,456 ↓ · 83 ♡

Qwen3-VL-30B-A3B-Instruct-GGUF

TensorBlock's GGUF conversion of Qwen3-VL-30B-A3B, Alibaba's 30B mixture-of-experts vision-language model activating 3B parameters per token. The GGUF packaging enables llama.cpp-compatible multimodal inference on consumer hardware.

388,024 ↓ · 20 ♡

gemma-4-31B-it-AWQ

gemma-4-31B-it-AWQ is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. Prebuilt AWQ weights make local and edge inference of gemma-4-31B-it-AWQ straightforward. The Apache 2.0 license keeps gemma-4-31B-it-AWQ unrestricted for commercial reuse. gemma-4-31B-it-AWQ is community-maintained, so track upstream changes and pin a known-good revision.

387,019 ↓ · 14 ♡

LocateAnything-3B

LocateAnything-3B is a 3-billion-parameter vision-language model from NVIDIA that performs open-vocabulary object grounding and localization via natural language queries. It is fine-tuned from Qwen2.5-3B-Instruct using NVIDIA's Eagle visual encoder framework and targets conversational grounding workflows. The model supports referring expression comprehension and visual question answering with spatial outputs.

384,700 ↓ · 2,904 ♡

Qwen3.6-35B-A3B-MLX-8bit

Qwen3.6-35B-A3B-MLX-8bit is an 8-bit MLX quantization of Qwen's 35B-parameter Mixture-of-Experts model (qwen3_5_moe), where only 3B parameters are active per forward pass. This makes it feasible to run a nominally large MoE model on Apple Silicon with reduced memory pressure.

384,097 ↓ · 2 ♡

gemma-3n-E2B-it

Gemma-3n-E2B-it is Google's instruction-tuned 2B edge model from the Gemma-3n family, combining image, audio, video, and text understanding in a single model. The 'n' suffix indicates the next-generation architecture with per-layer embeddings for efficiency. Gemma license applies — allows research and commercial use with restrictions.

381,993 ↓ · 320 ♡

MiniCPM-V-4_5

OpenBMB's MiniCPM-V-4.5 is an efficient multimodal vision-language model supporting single and multi-image inputs as well as video. With 1,094 likes it is one of the most community-validated efficient VL models, excelling in OCR, chart understanding, and visual question answering.

381,681 ↓ · 1,098 ♡

typhoon-ocr-3b

OpenTyphoon's 3B vision-language model specialized for OCR and document understanding across multiple languages. Built on Qwen2.5-VL, it is fine-tuned on multilingual document data to extract structured text from diverse document types including forms, tables, and mixed-language pages.

381,029 ↓ · 9 ♡

olmOCR-2-7B-1025-FP8

olmOCR-2-7B-1025-FP8 is AllenAI's FP8-quantized vision-language model for optical character recognition and document understanding, fine-tuned from Qwen2.5-VL-7B. It is optimized for extracting text from PDFs, research papers, and complex document layouts including tables, equations, and multi-column formats. The FP8 quantization allows deployment on a single A100 with reduced memory footprint.

380,889 ↓ · 253 ♡

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2

As a qwen-based large model, Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2 focuses on vision-language understanding. The Apache 2.0 license keeps Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2 unrestricted for commercial reuse. Weighing in near 27000M parameters, Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2 trades some ceiling for cheaper, faster inference. Before relying on Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2, reproduce its key numbers on representative inputs.

380,877 ↓ · 122 ♡

tiny-Qwen3_5ForConditionalGeneration-NoThink

tiny-Qwen3_5ForConditionalGeneration-NoThink is a deliberately minimal Qwen3.5 multimodal checkpoint created by the HuggingFace TRL team for unit-testing training code. NoThink indicates reasoning steps are disabled. High download counts reflect CI automation rather than real inference use.

380,869 ↓ · 0 ♡

dots.mocr

dots.mocr is RedNote's multimodal OCR model based on a custom Transformer architecture, designed for high-accuracy text extraction from documents including complex layouts, tables, formulas, and mixed Chinese-English content. It goes beyond standard OCR by understanding document structure, making it suitable for parsing invoices, forms, and academic papers in both Chinese and English.

379,856 ↓ · 141 ♡

tiny-Gemma3ForConditionalGeneration

A minimal HuggingFace Transformers library stub model implementing Gemma3's conditional generation architecture. Used internally for unit testing, CI pipelines, and rapid framework compatibility checks — not intended for any end-user generation task.

379,284 ↓ · 0 ♡

Qwen3.6-35B-A3B-MLX-6bit

Qwen3.6-35B-A3B-MLX-6bit is a 6-bit MLX quantization of Qwen/Qwen3.6-35B-A3B, a mixture-of-experts multimodal model targeting Apple Silicon via the MLX framework. With 35B total parameters and roughly 3B active per forward pass, it offers significantly reduced memory footprint compared to the full-precision base.

378,377 ↓ · 3 ♡

InternVL3-1B

InternVL3-1B is the 1B-parameter lightweight entry in Shanghai AI Lab's InternVL3 multimodal series, supporting image understanding and multilingual text generation. At this size it targets memory-constrained environments where larger VLMs are impractical.

377,634 ↓ · 84 ♡

Qwen2.5-VL-3B-Instruct-AWQ

Alibaba's AWQ 4-bit quantization of Qwen2.5-VL-3B-Instruct, the compact multimodal member of the Qwen2.5-VL family. AWQ compression enables a 3B VL model to run on very modest GPU hardware while maintaining near-BF16 visual understanding quality.

377,158 ↓ · 66 ♡

Ovis2.5-9B

AIDC-AI's Ovis2.5-9B is a 9B multimodal large language model supporting image-text-to-text tasks, with strong performance on charts, diagrams, and structured visual content. With 308 likes it has solid community validation as a capable mid-size VL model.

376,869 ↓ · 308 ♡

ChatRex-7B

ChatRex-7B is IDEA Research's 7B visual language model with an integrated Universal Proposal Network (UPN) for region-level grounding and object detection. Published in arxiv:2411.18363, it combines language understanding with fine-grained spatial reasoning over image regions. The base model is derived from LAION CLIP-convnext, indicating strong vision encoder pretraining.

375,574 ↓ · 14 ♡

Qwopus3.6-27B-v1-preview-GGUF

Qwopus3.6-27B is a GGUF-quantised preview fine-tune of Qwen3.6-27B positioned as a 'Claude Opus-style' reasoning and instruction model — the name blends Qwen and Opus. It targets advanced instruction following, multilingual reasoning, and multimodal (vision-language) tasks. As a v1 preview, evaluation is still community-driven and production use should be preceded by task-specific benchmarking.

375,239 ↓ · 125 ♡

Qwen3-VL-8B-Thinking

Qwen3-VL-8B-Thinking is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. The Apache 2.0 license keeps Qwen3-VL-8B-Thinking unrestricted for commercial reuse. Weighing in near 8000M parameters, Qwen3-VL-8B-Thinking trades some ceiling for cheaper, faster inference. Evaluate Qwen3-VL-8B-Thinking on your own data before trusting it in production.

374,948 ↓ · 210 ♡

tiny-Qwen3_5MoeForConditionalGeneration-3.6

tiny-Qwen3_5MoeForConditionalGeneration-3.6 is a minimal Qwen3.5 MoE multimodal checkpoint maintained by the HuggingFace TRL team for library testing purposes. The 3.6 suffix references the Qwen3.6 MoE variant shape. Like other TRL testing models, it is not intended for inference.

374,143 ↓ · 0 ♡

Qwopus3.6-27B-Coder-NVFP4

Bartowski's NVFP4-quantized version of Qwopus3.6-27B-Coder, a community coding merge of Qwen3.6-27B. NVFP4 targets Blackwell-generation GPU inference for maximally low memory use in local code generation.

373,463 ↓ · 3 ♡

InternVL3-78B-AWQ

AWQ INT4 quantization of InternVL3-78B, a large multimodal model from Shanghai AI Lab's OpenGVLab. InternVL3 covers image, video, and text understanding tasks. AWQ targets inference efficiency while preserving most of the base model's accuracy on VQA and captioning benchmarks. License is non-standard — review the model card before commercial use.

371,006 ↓ · 11 ♡

NVIDIA-Nemotron-Parse-v1.1

As a nemotron-based open-weight model, NVIDIA-Nemotron-Parse-v1.1 focuses on vision-language understanding. NVIDIA-Nemotron-Parse-v1.1 lists a non-standard license, so confirm permissions before deployment. Before relying on NVIDIA-Nemotron-Parse-v1.1, reproduce its key numbers on representative inputs.

369,745 ↓ · 171 ♡

GLM-4.1V-9B-Thinking

GLM-4.1V-9B-Thinking is Zhipu AI's 9B-parameter vision-language model with an integrated chain-of-thought reasoning module. The 'Thinking' variant explicitly generates internal reasoning steps before producing final answers, improving performance on complex visual question answering and multi-step visual reasoning tasks. It supports English and Chinese natively.

367,962 ↓ · 784 ♡

gemma-4-31B-it-MLX-8bit

An 8-bit MLX quantization of Google's Gemma 4 31B instruct model, prepared by the LM Studio community for Apple Silicon local inference. Gemma 4 31B is a dense instruction-tuned model targeting the mid-to-high capability tier.

367,599 ↓ · 2 ♡

gemma-4-31B-it-unsloth-bnb-4bit

As a gemma-based large model, gemma-4-31B-it-unsloth-bnb-4bit focuses on vision-language understanding. The Apache 2.0 license keeps gemma-4-31B-it-unsloth-bnb-4bit unrestricted for commercial reuse. Weighing in near 31000M parameters, gemma-4-31B-it-unsloth-bnb-4bit trades some ceiling for cheaper, faster inference. Before relying on gemma-4-31B-it-unsloth-bnb-4bit, reproduce its key numbers on representative inputs.

366,723 ↓ · 22 ♡

gemma-4-31B-it-FP8-dynamic

RedHatAI's FP8 dynamic quantization of Google's Gemma-4-31B-it using llm-compressor and compressed-tensors, compatible with vLLM. Dynamic FP8 applies per-tensor quantization at runtime for activations, reducing memory without a calibration dataset.

365,617 ↓ · 23 ♡

Qwen3.5-9B-AWQ-4bit

AWQ 4-bit quantization of Qwen3.5-9B, a dense image-text-to-text model. At 9B parameters with AWQ INT4, inference requires roughly 6-8 GB VRAM, placing it within reach of RTX 3080/4070-class cards. compressed-tensors format is vLLM-native.

363,295 ↓ · 36 ♡

gemma-4-31B-it-AWQ-4bit

gemma-4-31B-it-AWQ-4bit is an openly licensed vision-language understanding model in the gemma family. Prebuilt AWQ/4BIT weights make local and edge inference of gemma-4-31B-it-AWQ-4bit straightforward. It is a fine-tune of gemma-4-31b-it, inheriting that base model's general competence. Evaluate gemma-4-31B-it-AWQ-4bit on your own data before trusting it in production.

359,340 ↓ · 51 ♡

dots.ocr

dots.ocr is an image-text-to-text model specializing in optical character recognition with layout understanding, table extraction, and mathematical formula parsing. With 1,318 likes and 452K downloads it has strong community adoption for structured document digitization.

358,798 ↓ · 1,321 ♡

Qwen3.6-27B-Uncensored-HauhauCS-Aggressive

An 'aggressive' uncensored abliterated GGUF variant of Qwen3.6-27B, with safety refusal mechanisms removed via abliteration. Available in imatrix-calibrated GGUF quantizations. Safety removals affect the model's ability to decline harmful requests — this is a community fine-tune without safety evaluation.

355,031 ↓ · 569 ♡

Qwen3.5-27B-GPTQ-Int4

Official Alibaba GPTQ INT4 quantization of Qwen3.5-27B, a dense multimodal model for image and text tasks. GPTQ INT4 reduces memory to approximately 15-18 GB, making the model accessible on A100 or RTX 4090-class hardware. Apache-2.0 licensed.

354,659 ↓ · 55 ♡

Step3-VL-10B

Step3-VL-10B is StepFun's 10B-parameter vision-language model with a custom transformer architecture (step_robotics). It targets multimodal understanding tasks including image captioning, visual QA, and document reading. The model uses safetensors weights with custom inference code and is positioned as a mid-size VLM in StepFun's model family.

351,926 ↓ · 411 ♡

Intern-S1-Pro

Shanghai AI Lab's multimodal reasoning model from the InternLM family, targeting complex image-text reasoning tasks with a focus on chain-of-thought and multi-step problem solving over visual inputs. Two recent papers document the training approach (arxiv:2603.25040, 2508.15763). Apache 2.0 licensed.

351,910 ↓ · 279 ♡

gemma-4-31B-it-NVFP4

gemma-4-31B-it-NVFP4 is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. Weighing in near 31000M parameters, gemma-4-31B-it-NVFP4 trades some ceiling for cheaper, faster inference. The Apache 2.0 license keeps gemma-4-31B-it-NVFP4 unrestricted for commercial reuse. Treat gemma-4-31B-it-NVFP4's published metrics as a starting point and validate against your workload.

350,836 ↓ · 57 ♡

kanana-1.5-v-3b-instruct

Kanana-1.5-V is Kakao's 3B vision-language instruct model, part of their Kanana model family targeting Korean-English bilingual multimodal tasks. Optimized for practical deployment at 3B parameters while maintaining decent visual understanding.

350,061 ↓ · 56 ♡

gemma-3-27b-it-quantized.w4a16

As a gemma3-based large model, gemma-3-27b-it-quantized.w4a16 focuses on vision-language understanding. Weighing in near 27000M parameters, gemma-3-27b-it-quantized.w4a16 trades some ceiling for cheaper, faster inference. gemma-3-27b-it-quantized.w4a16 is subject to Gemma terms, so confirm licensing before commercial use. Read gemma-3-27b-it-quantized.w4a16's card for hardware requirements and licensing fine print before deploying.

349,612 ↓ · 13 ♡

Qwen3.5-122B-A10B-AWQ-4bit

Qwen3.5-122B-A10B-AWQ-4bit is an openly licensed vision-language understanding model in the qwen3 family. Qwen3.5-122B-A10B-AWQ-4bit is Apache 2.0-licensed, clearing it for closed-source and paid products. Prebuilt AWQ/4BIT weights make local and edge inference of Qwen3.5-122B-A10B-AWQ-4bit straightforward. Evaluate Qwen3.5-122B-A10B-AWQ-4bit on your own data before trusting it in production.

349,099 ↓ · 40 ♡

InternVL3-8B-AWQ

InternVL3-8B-AWQ is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. InternVL3-8B-AWQ is multilingual by design rather than English-only. Weighing in near 8000M parameters, InternVL3-8B-AWQ trades some ceiling for cheaper, faster inference. Evaluate InternVL3-8B-AWQ on your own data before trusting it in production.

348,390 ↓ · 8 ♡

Phi-3.5-vision-instruct-int8-ov

Phi-3.5-vision-instruct-int8-ov is an INT8-quantized version of Microsoft's Phi-3.5-Vision-Instruct, converted to OpenVINO IR format for CPU and Intel GPU inference. It targets edge and on-premise deployments where NVIDIA GPUs are unavailable or cost-prohibitive. The MIT license makes it suitable for commercial use, though the OpenVINO format limits portability to non-Intel runtimes.

346,837 ↓ · 2 ♡

Qwen3.5-9B-DeepSeek-V4-Flash-GGUF

A GGUF-quantised Qwen3.5-9B fine-tuned with DeepSeek V4 Flash distillation — the model has been trained on reasoning traces from a larger DeepSeek teacher to improve chain-of-thought quality at 9B scale. It targets multilingual reasoning with long-context CoT traces, supporting English, Chinese, Korean, Japanese, Spanish, and Russian. The GGUF format enables llama.cpp local inference.

346,746 ↓ · 239 ♡

Qwen3-VL-32B-Thinking-FP8

Qwen3-VL-32B-Thinking-FP8 is a 32B FP8-quantized Qwen3 vision-language model with extended reasoning ('Thinking') mode, enabling multi-step chain-of-thought for complex visual analysis tasks. FP8 quantization allows it to run on a single 80GB GPU rather than requiring multi-GPU setup for the full BF16 model. The Thinking mode produces visible reasoning traces before the final answer, improving accuracy on math, logic, and diagram interpretation.

346,267 ↓ · 26 ♡

paligemma-3b-pt-224

PaliGemma 3B pretrained at 224×224 resolution is Google's compact vision-language model checkpoint before instruction fine-tuning. The PT (pretrained) variant is intended as a foundation for task-specific fine-tuning rather than direct deployment.

345,806 ↓ · 484 ♡

Idefics3-8B-Llama3

Idefics3-8B-Llama3 is HuggingFace's open multimodal model combining a SigLIP vision encoder with a Llama 3 8B language backbone. It is designed to follow instructions over interleaved image-text inputs and was released alongside training infrastructure on the HuggingFace Hub. The model is positioned as a fully open (weights + training code + datasets) alternative to commercial VLMs.

345,497 ↓ · 304 ♡

Qwen3.6-27B-GPTQ-Pro-4bit

Community GPTQ quantization of Qwen3.6-27B at 4-bit using the Marlin kernel, calibrated with a 'Pro' dataset and packaged with FOEM (Fused Operation Efficient Marlin) for throughput on compatible GPUs. This is a multimodal (image-text) model.

344,125 ↓ · 39 ♡

Qwen3.5-35B-A3B-AWQ

An AWQ 4-bit quantization of Qwen3.5-35B-A3B (a 35B MoE with 3B active parameters) by QuantTrio, enabling memory-efficient inference on single high-VRAM GPUs. The MoE architecture means the 4-bit quantization applies to all expert weights rather than a dense 35B weight matrix.

343,576 ↓ · 18 ♡

Qwopus3.6-27B-v2-MTP-GGUF

Built for vision-language understanding, Qwopus3.6-27B-v2-MTP-GGUF is a model with publicly available weights. Qwopus3.6-27B-v2-MTP-GGUF is Apache 2.0-licensed, clearing it for closed-source and paid products. At about 27000M parameters, Qwopus3.6-27B-v2-MTP-GGUF sits in the large tier, which sets its memory and latency budget. Before relying on Qwopus3.6-27B-v2-MTP-GGUF, reproduce its key numbers on representative inputs.

342,530 ↓ · 340 ♡

SmolVLM2-2.2B-Instruct

Built for vision-language understanding, SmolVLM2-2.2B-Instruct is a model with publicly available weights. The weights start from smolvlm-instruct and specialize it for the target task. At about 2200M parameters, SmolVLM2-2.2B-Instruct sits in the mid-sized tier, which sets its memory and latency budget. Before relying on SmolVLM2-2.2B-Instruct, reproduce its key numbers on representative inputs.

342,056 ↓ · 324 ♡

Qwen3.5-397B-A17B-AWQ-4bit

As a qwen3-based frontier-scale model, Qwen3.5-397B-A17B-AWQ-4bit focuses on vision-language understanding. Weighing in near 397000M parameters, Qwen3.5-397B-A17B-AWQ-4bit trades some ceiling for cheaper, faster inference. AWQ builds of Qwen3.5-397B-A17B-AWQ-4bit are published alongside the full checkpoint for low-memory serving. Before relying on Qwen3.5-397B-A17B-AWQ-4bit, reproduce its key numbers on representative inputs.

340,681 ↓ · 3 ♡

InternVL3-1B-hf

InternVL3-1B-hf targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. Licensing for InternVL3-1B-hf is unspecified or custom — clear it before commercial use. InternVL3-1B-hf's 1000M-parameter size keeps hosting requirements modest relative to frontier models. InternVL3-1B-hf is community-maintained, so track upstream changes and pin a known-good revision.

338,460 ↓ · 10 ♡

Qwen3.6-27B-Heretic-Uncensored-FINETUNE-NEO-CODE-Di-IMatrix-MAX-GGUF

A 27B-parameter GGUF quantization of Qwen3.6 fine-tuned for creative writing, fiction, and code, with abliterated safety filters via the Heretic series. The imatrix quantization preserves perplexity better than naive integer quantization at the same bit-width.

338,299 ↓ · 366 ♡

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF is an openly licensed vision-language understanding model in the qwen family. At about 27000M parameters, Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF sits in the large tier, which sets its memory and latency budget. Prebuilt GGUF weights make local and edge inference of Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF straightforward. Treat Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-GGUF's published metrics as a starting point and validate against your workload.

332,853 ↓ · 656 ♡

gemma-3-27b-it-abliterated

An abliterated version of Google's Gemma-3-27B-IT, with safety refusal mechanisms removed by mlabonne using directional activation manipulation. Gemma license applies to the underlying weights. The abliteration removes content restrictions while preserving the model's multimodal instruction-following capability.

324,867 ↓ · 322 ♡

gemma-3n-E4B-it

Gemma 3n E4B Instruct repackaged by Unsloth for efficient local fine-tuning and inference. Gemma 3n is Google's on-device model family designed for mobile and edge hardware; E4B uses per-layer selective parameter activation to run with approximately 4B effective parameters while having a larger total capacity. Unsloth's repackage enables QLoRA fine-tuning of this model on consumer GPUs.

318,961 ↓ · 10 ♡

Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled

Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled is a large checkpoint for vision-language understanding, distributed on the HuggingFace Hub. Weighing in near 9000M parameters, Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled trades some ceiling for cheaper, faster inference. Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled is multilingual by design rather than English-only. Evaluate Qwen3.5-9B-Claude-4.6-Opus-Reasoning-Distilled on your own data before trusting it in production.

318,006 ↓ · 60 ♡

translategemma-4b-it

TranslateGemma-4b-it is Google's Gemma 3-based 4B instruction-tuned model fine-tuned specifically for translation tasks. Unlike generic multilingual LLMs, it was trained with translation as a primary objective, producing more accurate and fluent translations than prompting a general-purpose model. It uses the standard HuggingFace transformers interface for translation inference.

317,340 ↓ · 780 ♡

Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive

Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive is a qwen3-based open-weight model aimed at vision-language understanding. GGUF builds of Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive are published alongside the full checkpoint for low-memory serving. Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive's 35000M-parameter size keeps hosting requirements modest relative to frontier models. Check the Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive model card for benchmarks and intended use before adopting it.

316,109 ↓ · 1,407 ♡

Qianfan-OCR

Qianfan-OCR is Baidu's vision-language model specialized for optical character recognition and document intelligence, supporting multilingual text extraction from images. It combines a vision encoder with a language model for scene text understanding beyond simple character recognition. Apache-2.0 licensed with published benchmark results.

313,490 ↓ · 1,176 ♡

Qwopus3.6-35B-A3B-v1-GGUF

Qwopus3.6-35B-A3B-v1-GGUF is an openly licensed vision-language understanding model. Qwopus3.6-35B-A3B-v1-GGUF is multilingual by design rather than English-only. Qwopus3.6-35B-A3B-v1-GGUF is Apache 2.0-licensed, clearing it for closed-source and paid products. Qwopus3.6-35B-A3B-v1-GGUF is community-maintained, so track upstream changes and pin a known-good revision.

313,403 ↓ · 203 ♡

Qwen3.5-35B-A3B-GGUF

Qwen3.5-35B-A3B-GGUF is a qwen3-based open-weight model aimed at vision-language understanding. Qwen3.5-35B-A3B-GGUF's 35000M-parameter size keeps hosting requirements modest relative to frontier models. Permissive Apache 2.0 terms let Qwen3.5-35B-A3B-GGUF go straight into commercial pipelines. Qwen3.5-35B-A3B-GGUF ships without a hosted SLA, so budget for self-managed deployment and monitoring.

312,514 ↓ · 843 ♡

Kimi-K2.6-GGUF

Unsloth's GGUF conversion of Moonshot AI's Kimi K2.6 MoE model, enabling local inference via llama.cpp. Kimi K2 is a large MoE model from Moonshot AI notable for its strong reasoning performance at a competitive compute cost.

309,012 ↓ · 157 ♡

gemma-3n-E4B-it-MLX-bf16

gemma-3n-E4B-it-MLX-bf16 is a mid-sized checkpoint for vision-language understanding, distributed on the HuggingFace Hub. It is a fine-tune of gemma-3n-e4b-it, inheriting that base model's general competence. Weighing in near 4000M parameters, gemma-3n-E4B-it-MLX-bf16 trades some ceiling for cheaper, faster inference. Like most open checkpoints, gemma-3n-E4B-it-MLX-bf16 rewards a quick in-domain eval before commitment.

308,673 ↓ · 3 ♡

Qwen3-VL-235B-A22B-Instruct-FP8

Built for vision-language understanding, Qwen3-VL-235B-A22B-Instruct-FP8 is a qwen3-based model with publicly available weights. At about 235000M parameters, Qwen3-VL-235B-A22B-Instruct-FP8 sits in the frontier-scale tier, which sets its memory and latency budget. Qwen3-VL-235B-A22B-Instruct-FP8 is Apache 2.0-licensed, clearing it for closed-source and paid products. Qwen3-VL-235B-A22B-Instruct-FP8 ships without a hosted SLA, so budget for self-managed deployment and monitoring.

308,125 ↓ · 44 ♡

InternVL2_5-8B

InternVL2_5-8B is an internvl-based open-weight model aimed at vision-language understanding. Permissive MIT terms let InternVL2_5-8B go straight into commercial pipelines. Training spans multiple languages, so InternVL2_5-8B covers cross-lingual vision-language understanding from one checkpoint. InternVL2_5-8B ships without a hosted SLA, so budget for self-managed deployment and monitoring.

308,120 ↓ · 104 ♡

Qwen3.5-27B-AWQ

QuantTrio's AWQ 4-bit quantisation of Qwen3.5-27B, a multimodal image-text model at 27 billion parameters. This variant uses vLLM-compatible AWQ serialisation and targets teams running the 27B model on GPU servers with constrained memory. QuantTrio maintains several AWQ quantisations of Qwen family models with consistent quantisation settings.

307,462 ↓ · 43 ♡

gemma-3n-E4B-it-MLX-8bit

gemma-3n-E4B-it-MLX-8bit is an open-weight vision-language understanding model in the gemma family. At about 4000M parameters, gemma-3n-E4B-it-MLX-8bit sits in the mid-sized tier, which sets its memory and latency budget. Prebuilt MLX/8BIT weights make local and edge inference of gemma-3n-E4B-it-MLX-8bit straightforward. Like most open checkpoints, gemma-3n-E4B-it-MLX-8bit rewards a quick in-domain eval before commitment.

307,139 ↓ · 0 ♡

NVIDIA-Nemotron-Nano-12B-v2-VL-FP8

Built for vision-language understanding, NVIDIA-Nemotron-Nano-12B-v2-VL-FP8 is a nemotron-based model with publicly available weights. At about 12000M parameters, NVIDIA-Nemotron-Nano-12B-v2-VL-FP8 sits in the large tier, which sets its memory and latency budget. FP8 builds of NVIDIA-Nemotron-Nano-12B-v2-VL-FP8 are published alongside the full checkpoint for low-memory serving. NVIDIA-Nemotron-Nano-12B-v2-VL-FP8 ships without a hosted SLA, so budget for self-managed deployment and monitoring.

306,743 ↓ · 50 ♡

google_gemma-4-26B-A4B-it-GGUF

google_gemma-4-26B-A4B-it-GGUF is a gemma-based open-weight model aimed at vision-language understanding. GGUF builds of google_gemma-4-26B-A4B-it-GGUF are published alongside the full checkpoint for low-memory serving. Permissive Apache 2.0 terms let google_gemma-4-26B-A4B-it-GGUF go straight into commercial pipelines. google_gemma-4-26B-A4B-it-GGUF ships without a hosted SLA, so budget for self-managed deployment and monitoring.

304,592 ↓ · 113 ♡

Mistral-Small-3.2-24B-Instruct-2506-bnb-4bit

Built for vision-language understanding, Mistral-Small-3.2-24B-Instruct-2506-bnb-4bit is a mistral-based model with publicly available weights. Training spans multiple languages, so Mistral-Small-3.2-24B-Instruct-2506-bnb-4bit covers cross-lingual vision-language understanding from one checkpoint. Mistral-Small-3.2-24B-Instruct-2506-bnb-4bit is Apache 2.0-licensed, clearing it for closed-source and paid products. Mistral-Small-3.2-24B-Instruct-2506-bnb-4bit ships without a hosted SLA, so budget for self-managed deployment and monitoring.

304,064 ↓ · 10 ♡

Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GPTQ-int4

Built for vision-language understanding, Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GPTQ-int4 is a qwen3-based model with publicly available weights. At about 35000M parameters, Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GPTQ-int4 sits in the frontier-scale tier, which sets its memory and latency budget. Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GPTQ-int4 is Apache 2.0-licensed, clearing it for closed-source and paid products. Before relying on Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GPTQ-int4, reproduce its key numbers on representative inputs.

303,726 ↓ · 9 ♡

gemma-3n-E4B-it-MLX-6bit

gemma-3n-E4B-it-MLX-6bit is an open-weight vision-language understanding model in the gemma family. Distribution of gemma-3n-E4B-it-MLX-6bit is under Gemma, which is worth reading before you ship. Prebuilt MLX weights make local and edge inference of gemma-3n-E4B-it-MLX-6bit straightforward. Evaluate gemma-3n-E4B-it-MLX-6bit on your own data before trusting it in production.

303,662 ↓ · 0 ♡

Qwen3-VL-2B-Instruct-FP8

As a qwen3-based mid-sized model, Qwen3-VL-2B-Instruct-FP8 focuses on vision-language understanding. The Apache 2.0 license keeps Qwen3-VL-2B-Instruct-FP8 unrestricted for commercial reuse. FP8 builds of Qwen3-VL-2B-Instruct-FP8 are published alongside the full checkpoint for low-memory serving. Check the Qwen3-VL-2B-Instruct-FP8 model card for benchmarks and intended use before adopting it.

303,557 ↓ · 39 ♡

RolmOCR

RolmOCR is an openly licensed vision-language understanding model in the olmo family. RolmOCR is Apache 2.0-licensed, clearing it for closed-source and paid products. It is a fine-tune of qwen2.5-vl-7b-instruct, inheriting that base model's general competence. RolmOCR is community-maintained, so track upstream changes and pin a known-good revision.

302,994 ↓ · 586 ♡

gemma-3-27b-it-AWQ-INT4

gemma-3-27b-it-AWQ-INT4 targets vision-language understanding and is shipped as a large, self-hostable checkpoint. Prebuilt AWQ/INT4 weights make local and edge inference of gemma-3-27b-it-AWQ-INT4 straightforward. Permissive Apache 2.0 terms let gemma-3-27b-it-AWQ-INT4 go straight into commercial pipelines. Treat gemma-3-27b-it-AWQ-INT4's published metrics as a starting point and validate against your workload.

300,315 ↓ · 7 ♡

gemma-3-4b-it-qat-4bit

gemma-3-4b-it-qat-4bit targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. gemma-3-4b-it-qat-4bit is multilingual by design rather than English-only. gemma-3-4b-it-qat-4bit's 4000M-parameter size keeps hosting requirements modest relative to frontier models. Like most open checkpoints, gemma-3-4b-it-qat-4bit rewards a quick in-domain eval before commitment.

300,091 ↓ · 8 ♡

Gemma-4-E2B-Uncensored-HauhauCS-Aggressive

As a gemma-based mid-sized model, Gemma-4-E2B-Uncensored-HauhauCS-Aggressive focuses on vision-language understanding. Training spans multiple languages, so Gemma-4-E2B-Uncensored-HauhauCS-Aggressive covers cross-lingual vision-language understanding from one checkpoint. GGUF builds of Gemma-4-E2B-Uncensored-HauhauCS-Aggressive are published alongside the full checkpoint for low-memory serving. Check the Gemma-4-E2B-Uncensored-HauhauCS-Aggressive model card for benchmarks and intended use before adopting it.

299,361 ↓ · 165 ♡

gemma-4-31B-it-MLX-4bit

gemma-4-31B-it-MLX-4bit targets vision-language understanding and is shipped as a large, self-hostable checkpoint. gemma-4-31B-it-MLX-4bit's 31000M-parameter size keeps hosting requirements modest relative to frontier models. Prebuilt MLX/4BIT weights make local and edge inference of gemma-4-31B-it-MLX-4bit straightforward. Like most open checkpoints, gemma-4-31B-it-MLX-4bit rewards a quick in-domain eval before commitment.

298,419 ↓ · 1 ♡

Qwen3.5-0.8B-GGUF

Unsloth's GGUF conversion of Qwen3.5-0.8B, the smallest model in the Qwen3.5 series. At 0.8B parameters, it targets extremely constrained inference environments — Raspberry Pi, microcontrollers with GGUF support, or embedding in applications.

297,370 ↓ · 178 ♡

Qwen3.5-2B-GGUF

Qwen3.5-2B-GGUF is a mid-sized checkpoint for vision-language understanding, distributed on the HuggingFace Hub. Weighing in near 2000M parameters, Qwen3.5-2B-GGUF trades some ceiling for cheaper, faster inference. Prebuilt GGUF weights make local and edge inference of Qwen3.5-2B-GGUF straightforward. Treat Qwen3.5-2B-GGUF's published metrics as a starting point and validate against your workload.

296,345 ↓ · 100 ♡

Qwopus3.5-9B-v3

Built for vision-language understanding, Qwopus3.5-9B-v3 is a qwen-based model with publicly available weights. Training spans multiple languages, so Qwopus3.5-9B-v3 covers cross-lingual vision-language understanding from one checkpoint. Qwopus3.5-9B-v3 is Apache 2.0-licensed, clearing it for closed-source and paid products. Qwopus3.5-9B-v3 ships without a hosted SLA, so budget for self-managed deployment and monitoring.

294,775 ↓ · 88 ♡

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF is a qwen-based open-weight model aimed at vision-language understanding. Training spans multiple languages, so Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF covers cross-lingual vision-language understanding from one checkpoint. Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF's 27000M-parameter size keeps hosting requirements modest relative to frontier models. Before relying on Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF, reproduce its key numbers on representative inputs.

292,155 ↓ · 601 ♡

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled

Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled targets vision-language understanding and is shipped as a large, self-hostable checkpoint. It is a fine-tune of qwen3.5-27b, inheriting that base model's general competence. Permissive Apache 2.0 terms let Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled go straight into commercial pipelines. Evaluate Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled on your own data before trusting it in production.

290,793 ↓ · 2,814 ♡

gemma-3n-E4B-it-MLX-4bit

gemma-3n-E4B-it-MLX-4bit targets vision-language understanding and is shipped as a mid-sized, self-hostable checkpoint. Prebuilt MLX/4BIT weights make local and edge inference of gemma-3n-E4B-it-MLX-4bit straightforward. gemma-3n-E4B-it-MLX-4bit's 4000M-parameter size keeps hosting requirements modest relative to frontier models. Treat gemma-3n-E4B-it-MLX-4bit's published metrics as a starting point and validate against your workload.

289,004 ↓ · 2 ♡

Qwen3.5-9B-FP8

Qwen3.5-9B-FP8 is a qwen3-based open-weight model aimed at vision-language understanding. Permissive Apache 2.0 terms let Qwen3.5-9B-FP8 go straight into commercial pipelines. Qwen3.5-9B-FP8's 9000M-parameter size keeps hosting requirements modest relative to frontier models. Qwen3.5-9B-FP8 ships without a hosted SLA, so budget for self-managed deployment and monitoring.

287,785 ↓ · 10 ♡

google_gemma-4-31B-it-GGUF

google_gemma-4-31B-it-GGUF is a gemma-based open-weight model aimed at vision-language understanding. Permissive Apache 2.0 terms let google_gemma-4-31B-it-GGUF go straight into commercial pipelines. google_gemma-4-31B-it-GGUF's 31000M-parameter size keeps hosting requirements modest relative to frontier models. google_gemma-4-31B-it-GGUF ships without a hosted SLA, so budget for self-managed deployment and monitoring.

285,205 ↓ · 62 ♡

llava-v1.5-7b

LLaVA 1.5 7B is Haotian Liu et al.'s multimodal instruction-following model combining a CLIP vision encoder with a Vicuna-7B language model. At 7B, it was one of the strongest open VLMs at its release and remains a common fine-tuning starting point.

235,049 ↓ · 555 ♡