{"object":"list","data":[{"id":"voxtral-mini-4b-realtime","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":[],"aliases":["stt"],"model_type":"asr","capabilities":{},"context_length":32768,"model_full_name":"Voxtral Realtime 4B","model_icon_url":"/model-icons/mistral.svg","description":"Real-time multilingual speech-to-text","use_when":"Mistral's 4B streaming ASR built on a causal audio encoder, delivering sub-500ms configurable-delay transcription over a websocket across 13 languages. Audio-in, text-out only. Use it for live captions, real-time transcription, meeting notes, and voice assistants.","cost":{"input_per_m_token":0.1,"output_per_m_token":0.3},"open_weights":true},{"id":"nvidia-gemma-4-31b-it-nvfp4","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty","tools","tool_choice","response_format","image_url"],"aliases":["fast"],"model_type":"vision-chat","capabilities":{"tools":true,"response_format":true,"image_input":true},"context_length":131072,"model_full_name":"Gemma 4 31B","model_icon_url":"/model-icons/google.svg","description":"Non-thinking vision chat with tool use","use_when":"NVIDIA's NVFP4 4-bit build of Google's Gemma 4 31B IT for B200-class GPUs, holding near-BF16 quality (85.35 vs 85.80 on the vendor benchmark). Non-thinking vision chat with tool-calling and image input (up to 4 images); Apache-2.0, 128K context.","cost":{"input_per_m_token":0.5,"output_per_m_token":1.1},"open_weights":true},{"id":"chatterbox-turbo","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":[],"aliases":["tts"],"model_type":"tts","capabilities":{},"model_full_name":"Chatterbox Turbo 350M","model_icon_url":"/model-icons/resemble.svg","description":"Low-latency streaming text-to-speech","use_when":"A 350M streaming text-to-speech model with a single-step distilled decoder for real-time, low-latency audio out. Clones a voice from a ~10s reference clip and takes expressiveness controls. Every output WAV carries a Perth neural watermark that can't be disabled.","cost":{"input_per_m_token":0.1,"output_per_m_token":0.3},"open_weights":true},{"id":"olmocr-2-7b-1025-fp8","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty"],"aliases":["ocr"],"model_type":"ocr-vision","capabilities":{},"context_length":16384,"model_full_name":"olmOCR 2 7B","model_icon_url":"/model-icons/allen-ai.svg","description":"Document OCR for scanned pages and PDFs","use_when":"A Qwen2.5-VL-7B fine-tune purpose-built for document OCR, not general vision-chat. FP8-quantized to fit one GPU. Extracts text from scanned pages and PDFs and emits YAML metadata (language, rotation, table/diagram flags) alongside the text; 16K context.","cost":{"input_per_m_token":0.1,"output_per_m_token":0.3},"open_weights":true},{"id":"qwen3-vl-32b-thinking-fp8","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty","reasoning","include_reasoning","tools","tool_choice","response_format","image_url"],"aliases":["vision"],"model_type":"vision-chat","capabilities":{"tools":true,"response_format":true,"image_input":true},"context_length":32768,"model_full_name":"Qwen 3 VL Thinking 32B","model_icon_url":"/model-icons/qwen.svg","description":"Vision reasoning with visible thinking","use_when":"A 32B FP8 vision-language model that reasons over interleaved image and text with an explicit, visible chain-of-thought. Thinking-only, so it always emits a reasoning trace. Reach for it on visual Q&A and multimodal tasks that benefit from step-by-step reasoning; deployed at 32K context.","cost":{"input_per_m_token":0.5,"output_per_m_token":1.1},"open_weights":true},{"id":"mistral-medium-3-5-128b-nvfp4","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty","reasoning","include_reasoning","tools","tool_choice","response_format","reasoning_effort"],"aliases":[],"model_type":"chat","capabilities":{"tools":true,"response_format":true,"reasoning_effort":true},"context_length":131072,"model_full_name":"Mistral Medium 3.5 128B","model_icon_url":"/model-icons/mistral.svg","description":"Balanced dense chat with optional thinking","use_when":"The NVFP4 build of Mistral Medium 3.5, a dense 128B model (FP4 MLP weights, FP8 attention/KV) tuned for balanced speed and quality on everyday chat, drafting, and code. Supports tool calls, structured output, and an optional step-by-step thinking mode. Text-only at 128K context.","cost":{"input_per_m_token":1.7,"output_per_m_token":7.5},"open_weights":true},{"id":"nvidia-kimi-k2-6-nvfp4","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty","reasoning","include_reasoning","tools","tool_choice","response_format","reasoning_effort","image_url"],"aliases":["coding","reasoning"],"model_type":"vision-chat","capabilities":{"tools":true,"response_format":true,"reasoning_effort":true,"image_input":true},"context_length":262144,"model_full_name":"Kimi K2.6 1T-32B","model_icon_url":"/model-icons/moonshot.svg","description":"1T MoE vision chat, NVFP4 for B200","use_when":"NVIDIA's NVFP4 repack of Moonshot's Kimi K2.6 for B200-class GPUs: a 1T-parameter MoE (32B active) with vision, 256K context, and agentic tool use. Unlike the base INT4 build, it surfaces its thinking as a separate reasoning trace. Best for multimodal, long-context, multi-step work on Blackwell.","cost":{"input_per_m_token":2.0,"output_per_m_token":5.0},"open_weights":true},{"id":"buffettbot-fp8-dynamic","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty"],"aliases":[],"model_type":"chat","capabilities":{},"context_length":32768,"model_full_name":"BuffettBot 32B","model_icon_url":"/model-icons/buffettbot.svg","description":"Investing prose in a Buffett-style voice","use_when":"Options IT's in-house 32B fine-tune for finance and investment writing, tuned toward a plain-spoken, value-investing (Buffett-style) voice. Reach for it to draft explainers, brainstorm content, and answer investment questions. Text-only chat, FP8-quantized, 32K context.","cost":{"input_per_m_token":0.1,"output_per_m_token":0.3},"open_weights":false},{"id":"qwen3-5-122b-a10b-nvfp4","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty","reasoning","include_reasoning","tools","tool_choice","response_format","reasoning_effort","image_url"],"aliases":["default"],"model_type":"vision-chat","capabilities":{"tools":true,"response_format":true,"reasoning_effort":true,"image_input":true},"context_length":262144,"model_full_name":"Qwen 3.5 122B-10B","model_icon_url":"/model-icons/qwen.svg","description":"Vision-capable coding with hybrid thinking","use_when":"The NVFP4 (Blackwell) build of Qwen 3.5's 122B MoE (10B active). Vision-capable, with hybrid thinking on by default, 256K context, and strong coding and agentic tool use. Reach for vision, long-context, or reasoning-heavy work.","cost":{"input_per_m_token":0.8,"output_per_m_token":3.0},"open_weights":true},{"id":"qwen3-reranker-4b","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty"],"aliases":["reranker"],"model_type":"reranker","capabilities":{},"context_length":8192,"model_full_name":"Qwen3 Reranker 4B","model_icon_url":"/model-icons/qwen.svg","description":"Re-ranks retrieval hits by relevance","use_when":"A 4B cross-encoder reranker that jointly scores each query-document pair and returns a calibrated [0,1] relevance score. Use it to re-rank vector-search candidates before they reach the model, sharpening RAG precision. Not a chat model; 8K input limit per pair.","cost":{"input_per_m_token":0.1,"output_per_m_token":0.0},"open_weights":true},{"id":"qwen3guard-gen-8b","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty"],"aliases":[],"model_type":"guard","capabilities":{},"context_length":32768,"model_full_name":"Qwen3Guard 8B","model_icon_url":"/model-icons/qwen.svg","description":"Flags unsafe prompts and responses","use_when":"The generative member of Alibaba's Qwen3Guard family: an 8B safety-moderation classifier that returns a Safe / Unsafe / Controversial verdict plus the violated categories across 119 languages, with a Refusal line for assistant responses. Powers the internal /v1/moderations gate; not a chat model.","cost":{"input_per_m_token":0.1,"output_per_m_token":0.1},"open_weights":true},{"id":"nemotron-3-super-120b-a12b-nvfp4","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":["max_tokens","temperature","top_p","top_k","stream","stop","seed","frequency_penalty","presence_penalty","reasoning","include_reasoning","tools","tool_choice","response_format","reasoning_effort"],"aliases":["summary"],"model_type":"chat","capabilities":{"tools":true,"response_format":true,"reasoning_effort":true},"context_length":262144,"model_full_name":"Nemotron 3 Super 120B-12B","model_icon_url":"/model-icons/nvidia.svg","description":"Reasoning and code with visible thinking","use_when":"A 120B-A12B MoE (Mamba-2 + Latent MoE) tuned for reasoning-heavy code, analysis, and agentic tool-calling with visible thinking and adjustable effort. This is the NVFP4 build, trained natively in 4-bit and sized to run on a single B200-class (Blackwell) GPU; text-only.","cost":{"input_per_m_token":0.8,"output_per_m_token":3.0},"open_weights":true},{"id":"kalm-embedding-gemma3-12b-2511","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":[],"aliases":["embeddings"],"model_type":"embeddings","capabilities":{},"context_length":8192,"model_full_name":"KaLM Embedding 12B","model_icon_url":"/model-icons/tencent.svg","description":"Matryoshka text embeddings for retrieval","use_when":"A 12B Gemma-3 dense embedding encoder producing 3840-dim vectors, truncatable via Matryoshka to 2048/1024/512 and smaller without re-encoding. Use for RAG retrieval, semantic search, and clustering; queries take an instruct prefix, documents none.","cost":{"input_per_m_token":0.1,"output_per_m_token":0.0},"open_weights":true},{"id":"cosmos3-super-text2image","object":"model","created":1785196410,"owned_by":"options-it","supported_parameters":[],"aliases":["t2i"],"model_type":"image-gen","capabilities":{},"model_full_name":"Cosmos 3 Super 64B","model_icon_url":"/model-icons/nvidia.svg","description":"Generates images from text prompts","use_when":"NVIDIA's 64B omni-modal diffusion model for text-to-image generation, served through the images endpoint. Reach for it to turn a text prompt into a picture. It produces images, not chat replies, and bills per generated image.","cost":{"input_per_m_token":0.0,"output_per_m_token":0.0,"image_per_generation":0.03},"open_weights":true}]}