Zero Vendor Lock-In • 13+ Providers • 60+ Tested Models

Supported LLM Providers & Models

Smoke Monkey Harness speaks standard OpenAI-compatible API schemas. Switch from NVIDIA Nemotron to Google Gemini, OpenAI o1/o3, DeepSeek R1, Groq LPUs, or 100% offline Ollama with a single configuration parameter.

13+
First-Class Providers
60+
Verified Tool Models
2,000+
Max tok/s (Cerebras WSE)
100%
Local / Offline Capable
Showing 13 of 13 providersStandard OpenAI-Compatible API Spec

NVIDIA Hosted NIM

Cloud / Hosted
Fast (~90 tok/s)128k context

High-throughput enterprise inference with reasoning token streaming and structured tool calling.

Tool CallingReasoning StreamAuto-Retry (5xx backoff)Production NIM
Environment Variable:NVIDIA_API_KEY
Tested Models (9):
nvidia/nemotron-3-super-120b-a12bDefault
Flagship agentic reasoning and tool execution on NVIDIA NIM.
nvidia/nemotron-3-ultra-550b-a55b
Ultra-scale MoE reasoning powerhouse for complex workflows.
nvidia/nemotron-3-nano-30b-a3b
Ultra-low latency, highly cost-effective agent reasoning.
nvidia/nemotron-3.5-lightning
Instant response times for fast verification and monitoring.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'nvidia',
  model: 'nvidia/nemotron-3-super-120b-a12b',
  apiKey: process.env.NVIDIA_API_KEY,
  tools: [readTool, writeTool, bashTool],
});

Google Gemini

Cloud API
Ultra-fast Flash / Deep Pro2M tokens context

Massive multimodal context with high-precision native function calling and ultra-fast Flash tiers.

2M ContextMultimodalNative Function CallingStreaming Deltas
Environment Variable:GEMINI_API_KEY
Tested Models (8):
gemini-2.5-proDefault
Google flagship reasoning model with massive multimodal capabilities.
gemini-2.5-flash
High-speed, cost-effective multimodal workhorse for agent loops.
gemini-2.5-flash-lite
Ultra-low latency inference for background cron monitors.
gemini-2.0-flash
Real-time speed with native agentic tool execution.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'gemini',
  model: 'gemini-2.5-pro',
  apiKey: process.env.GEMINI_API_KEY,
  tools: [searchWeb, analyzeCode],
});

OpenRouter (Universal Router)

Universal Router
Dynamic Multi-Provider Routing200k context

Access 200+ models with a single unified key, automatic failover, and price-to-performance routing.

200+ ModelsSingle KeyAuto-FallbackClaude + DeepSeek + GPT
Environment Variable:OPENROUTER_API_KEY
Tested Models (9):
anthropic/claude-3.7-sonnetDefault
State-of-the-art hybrid reasoning & coding intelligence from Anthropic.
anthropic/claude-3.5-sonnet
Gold standard coding agent model with superior tool reliability.
deepseek/deepseek-r1
High-intelligence open reasoning model with step-by-step thinking.
openai/gpt-4o
Flagship multimodal speed and tool calling performance.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'openrouter',
  model: 'anthropic/claude-3.7-sonnet',
  apiKey: process.env.OPENROUTER_API_KEY,
});

OpenAI

Cloud API
Fast (~80 tok/s)128k context

Industry benchmark for function calling, o-series deep reasoning, and high-speed GPT-4o execution.

Strict JSON Schemao1 / o3 Deep ReasoningParallel Tool CallingFlagship Reliability
Environment Variable:OPENAI_API_KEY
Tested Models (5):
gpt-4oDefault
High-speed flagship omni model for text, reasoning, and tools.
gpt-4o-mini
Fast, lightweight and cost-efficient for routine tasks.
o1
Full reasoning model designed for complex planning and coding.
o3-mini
Efficient reasoning model with adjustable reasoning effort.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'openai',
  model: 'gpt-4o',
  apiKey: process.env.OPENAI_API_KEY,
});

xAI — Grok

Cloud API
Fast (~75 tok/s)128k context

Real-time reasoning and code intelligence powered by xAI Grok with live web awareness.

Grok-2 ReasoningVision MultimodalReal-Time InsightsDirect API
Environment Variable:XAI_API_KEY
Tested Models (3):
grok-2-1212Default
Flagship reasoning and code intelligence from xAI.
grok-2-vision-1212
Multimodal vision and document analysis model.
grok-beta
Experimental fast inference preview.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'xai',
  model: 'grok-2-1212',
  apiKey: process.env.XAI_API_KEY,
});

DeepSeek AI Direct

Cloud API
Fast (~60 tok/s)64k context

Direct API access to DeepSeek V3 and R1 with full chain-of-thought token streaming at 1/10th market cost.

DeepSeek R1 CoTExtreme Cost EfficiencyOpen Weights ArchitectureCoding Excellence
Environment Variable:DEEPSEEK_API_KEY
Tested Models (2):
deepseek-chatDefault
General language and code powerhouse with top benchmark scores.
deepseek-reasoner
Full chain-of-thought reasoning for complex engineering decisions.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'deepseek',
  model: 'deepseek-chat',
  apiKey: process.env.DEEPSEEK_API_KEY,
});

Groq LPUs

Fast LPU Hardware
Extreme (500+ tok/s)128k context

Blazing fast inference (500+ tokens/second) on specialized Language Processing Units.

Sub-100ms First TokenLPU ArchitectureLlama 3.3 70BDeepSeek R1 Distill
Environment Variable:GROQ_API_KEY
Tested Models (4):
llama-3.3-70b-versatileDefault
Blazing fast inference speed (500+ tokens/sec).
llama-3.1-8b-instant
Sub-100ms response time for fast checks and scraping.
deepseek-r1-distill-llama-70b
High-speed reasoning model on specialized Groq LPUs.
mixtral-8x7b-32768
MoE high-throughput with 32k context.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'groq',
  model: 'llama-3.3-70b-versatile',
  apiKey: process.env.GROQ_API_KEY,
});

Cerebras Fast Inference

Wafer-Scale Engine
Insane (2,000+ tok/s)128k context

Extreme wafer-scale speed (2,000+ tokens/sec) on Cerebras CS-3 supercomputing systems.

2,000+ Tokens/SecWafer-Scale EngineInstant Tool VerificationZero Queuing
Environment Variable:CEREBRAS_API_KEY
Tested Models (2):
llama3.3-70bDefault
Extreme speed (2,000+ tokens/sec) on Cerebras Wafer-Scale Engine.
llama3.1-8b
Instantaneous sub-50ms latency for real-time monitoring.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'cerebras',
  model: 'llama3.3-70b',
  apiKey: process.env.CEREBRAS_API_KEY,
});

Mistral AI

Cloud API
Fast (~80 tok/s)128k context

Leading European frontier AI models with native function calling, 128k context, and Codestral.

Mistral Large 2Codestral 22BJSON ModeStructured Outputs
Environment Variable:MISTRAL_API_KEY
Tested Models (3):
mistral-large-latestDefault
Top-tier reasoning, 128k context, and benchmark-leading tool calling.
codestral-latest
Specialized 22B model for code completion and synthesis.
mistral-small-latest
Fast, responsive model for routine automation tasks.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'mistral',
  model: 'mistral-large-latest',
  apiKey: process.env.MISTRAL_API_KEY,
});

Qwen / Alibaba DashScope

Cloud API
Fast (~70 tok/s)128k context

State-of-the-art open coding and multilingual foundation models from Alibaba Cloud.

Qwen 2.5 CoderMultilingual 100+ langsHigh Math & Code BenchmarkDashScope Gateway
Environment Variable:QWEN_API_KEY
Tested Models (4):
qwen-maxDefault
Flagship Qwen foundation model with superior multilingual intelligence.
qwen2.5-coder-32b-instruct
Code-specialized model with high code generation precision.
qwen-plus
Balanced performance for general agentic tasks.
qwen-turbo
High-speed, cost-effective inference.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'qwen',
  model: 'qwen-max',
  apiKey: process.env.QWEN_API_KEY,
});

Together AI

Cloud API
Fast (~110 tok/s)128k context

High-performance hosted open source models on dedicated clusters with fast TTFT.

Llama 3.3 TurboDeepSeek R1 HostedOpen Weights ClusterDedicated Endpoints
Environment Variable:TOGETHER_API_KEY
Tested Models (3):
meta-llama/Llama-3.3-70B-Instruct-TurboDefault
High-speed open model on Together inference engine.
deepseek-ai/DeepSeek-R1
Open reasoning model hosted on Together infrastructure.
Qwen/Qwen2.5-Coder-32B-Instruct
Code-optimized open model.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'together',
  model: 'meta-llama/Llama-3.3-70B-Instruct-Turbo',
  apiKey: process.env.TOGETHER_API_KEY,
});

Ollama (Local & Offline)

Local Inference
Local Hardware Dependent32k - 128k context

100% offline, private, zero-telemetry local LLM inference running directly on Mac (Metal) or Linux/Windows (CUDA).

100% OfflineZero TelemetryApple Silicon Metal AccelKeyless Default
Environment Variable:None (Localhost:11434)
Tested Models (5):
qwen3:8bDefault
Recommended default: native function calling & high tool precision.
llama3.2:latest
Ultra-fast lightweight local model for routine tasks.
qwen2.5-coder:latest
Code-specialized local agent model.
deepseek-r1:latest
Local chain-of-thought reasoning without sending data to the cloud.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

// Runs out of the box with zero environment variables needed!
const agent = createAgent({
  provider: 'ollama',
  model: 'qwen3:8b',
  baseUrl: 'http://localhost:11434',
});
AI

OmniRoute / Gateway

Keyless Gateway
Fast Gateway32k context

Free tier keyless gateway and local model proxy with OpenAI-compatible routing.

Keyless Free TierCustom BaseUrl OverridesvLLM / LM Studio CompatibleMulti-tenant Routing
Environment Variable:OMNIROUTE_API_KEY (Optional)
Tested Models (3):
llama3.2:latestDefault
Default keyless endpoint model.
qwen2.5-coder:latest
Local/gateway code synthesis.
deepseek-r1:latest
Keyless reasoning preview.
Quickstart Snippet
import { createAgent } from 'smoke-monkey-harness';

const agent = createAgent({
  provider: 'omniroute',
  model: 'llama3.2:latest',
  baseUrl: 'https://api.getomni.app/openai/v1',
});

How Environment Key Resolution Works

The harness lets you specify API keys via environment variables, direct configuration strings, or an asynchronous key resolver callback. When deploying in multi-tenant environments, you can dynamically select per-user API keys at runtime:

import { createAgent } from 'smoke-monkey-harness';

// 1. Static Configuration
const agent = createAgent({
  provider: process.env.PROVIDER ?? 'nvidia',
  model: process.env.MODEL ?? 'nvidia/nemotron-3-super-120b-a12b',
  apiKey: process.env[`${(process.env.PROVIDER ?? 'nvidia').toUpperCase()}_API_KEY`],
});

// 2. Dynamic Per-User Key Resolver (Multi-Tenant SaaS)
const multiTenantAgent = createAgent({
  provider: 'openrouter',
  model: 'anthropic/claude-3.7-sonnet',
  apiKey: async (provider, userId) => {
    return await userKeyVault.getSecret(userId, provider);
  },
});

Explore More Architecture Resources