Supported LLM Providers & Models
Smoke Monkey Harness speaks standard OpenAI-compatible API schemas. Switch from NVIDIA Nemotron to Google Gemini, OpenAI o1/o3, DeepSeek R1, Groq LPUs, or 100% offline Ollama with a single configuration parameter.
NVIDIA Hosted NIM
High-throughput enterprise inference with reasoning token streaming and structured tool calling.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'nvidia',
model: 'nvidia/nemotron-3-super-120b-a12b',
apiKey: process.env.NVIDIA_API_KEY,
tools: [readTool, writeTool, bashTool],
});Google Gemini
Massive multimodal context with high-precision native function calling and ultra-fast Flash tiers.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'gemini',
model: 'gemini-2.5-pro',
apiKey: process.env.GEMINI_API_KEY,
tools: [searchWeb, analyzeCode],
});OpenRouter (Universal Router)
Access 200+ models with a single unified key, automatic failover, and price-to-performance routing.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'openrouter',
model: 'anthropic/claude-3.7-sonnet',
apiKey: process.env.OPENROUTER_API_KEY,
});OpenAI
Industry benchmark for function calling, o-series deep reasoning, and high-speed GPT-4o execution.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'openai',
model: 'gpt-4o',
apiKey: process.env.OPENAI_API_KEY,
});xAI — Grok
Real-time reasoning and code intelligence powered by xAI Grok with live web awareness.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'xai',
model: 'grok-2-1212',
apiKey: process.env.XAI_API_KEY,
});DeepSeek AI Direct
Direct API access to DeepSeek V3 and R1 with full chain-of-thought token streaming at 1/10th market cost.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'deepseek',
model: 'deepseek-chat',
apiKey: process.env.DEEPSEEK_API_KEY,
});Groq LPUs
Blazing fast inference (500+ tokens/second) on specialized Language Processing Units.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'groq',
model: 'llama-3.3-70b-versatile',
apiKey: process.env.GROQ_API_KEY,
});Cerebras Fast Inference
Extreme wafer-scale speed (2,000+ tokens/sec) on Cerebras CS-3 supercomputing systems.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'cerebras',
model: 'llama3.3-70b',
apiKey: process.env.CEREBRAS_API_KEY,
});Mistral AI
Leading European frontier AI models with native function calling, 128k context, and Codestral.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'mistral',
model: 'mistral-large-latest',
apiKey: process.env.MISTRAL_API_KEY,
});Qwen / Alibaba DashScope
State-of-the-art open coding and multilingual foundation models from Alibaba Cloud.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'qwen',
model: 'qwen-max',
apiKey: process.env.QWEN_API_KEY,
});Together AI
High-performance hosted open source models on dedicated clusters with fast TTFT.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'together',
model: 'meta-llama/Llama-3.3-70B-Instruct-Turbo',
apiKey: process.env.TOGETHER_API_KEY,
});Ollama (Local & Offline)
100% offline, private, zero-telemetry local LLM inference running directly on Mac (Metal) or Linux/Windows (CUDA).
import { createAgent } from 'smoke-monkey-harness';
// Runs out of the box with zero environment variables needed!
const agent = createAgent({
provider: 'ollama',
model: 'qwen3:8b',
baseUrl: 'http://localhost:11434',
});OmniRoute / Gateway
Free tier keyless gateway and local model proxy with OpenAI-compatible routing.
import { createAgent } from 'smoke-monkey-harness';
const agent = createAgent({
provider: 'omniroute',
model: 'llama3.2:latest',
baseUrl: 'https://api.getomni.app/openai/v1',
});How Environment Key Resolution Works
The harness lets you specify API keys via environment variables, direct configuration strings, or an asynchronous key resolver callback. When deploying in multi-tenant environments, you can dynamically select per-user API keys at runtime:
import { createAgent } from 'smoke-monkey-harness';
// 1. Static Configuration
const agent = createAgent({
provider: process.env.PROVIDER ?? 'nvidia',
model: process.env.MODEL ?? 'nvidia/nemotron-3-super-120b-a12b',
apiKey: process.env[`${(process.env.PROVIDER ?? 'nvidia').toUpperCase()}_API_KEY`],
});
// 2. Dynamic Per-User Key Resolver (Multi-Tenant SaaS)
const multiTenantAgent = createAgent({
provider: 'openrouter',
model: 'anthropic/claude-3.7-sonnet',
apiKey: async (provider, userId) => {
return await userKeyVault.getSecret(userId, provider);
},
});