agent, provider knobs under
harness_config.sdk_settings.local, and tool/sandbox policy at top level.
How to read this page
Decision rules per capability state — what YOU (the reader/agent) should do:
Parameter table columns: Parameter is the provider-side path the harness
maps for you (you set it via workflow fields or
harness_config.sdk_settings.<provider>
knobs — see Runtimes for which knobs exist). Default
applies when you say nothing. Allowed is exhaustive — values outside it
fail validation. Risk is deeda’s policy weight: high parameters
generally require elevated tool_policy/sandbox_profile and may trigger
human approval on hot-reload.
Core generation
Capabilities (4/4 rows usable):local.llamacpp.openai_compat(llm.complete) — vendor docslocal.llamacpp.server(llm.complete) — vendor docslocal.ollama.generate(llm.complete) — vendor docslocal.ollama.openai_compat(llm.complete) — vendor docs
llama-cpp-python
Capabilities (13/13 rows usable):local.llamacpp.python.create_chat_completion(llm.chat) — vendor docslocal.llamacpp.python.create_completion(llm.complete) — vendor docslocal.llamacpp.python.create_embedding(llm.embed) — vendor docslocal.llamacpp.python.eval_sample_generate(llm.streaming) — vendor docslocal.llamacpp.python.from_pretrained(provider.models) — vendor docslocal.llamacpp.python.llama_cache_state(provider.slots) — vendor docslocal.llamacpp.python.llama_class(provider.models) — vendor docslocal.llamacpp.python.llama_grammar(llm.structured_output) — vendor docslocal.llamacpp.python.logits_processor(llm.sampling) — vendor docslocal.llamacpp.python.save_load_state(provider.slots) — vendor docslocal.llamacpp.python.server_module(provider.lifecycle) — vendor docslocal.llamacpp.python.stopping_criteria(llm.sampling) — vendor docslocal.llamacpp.python.tokenize_detokenize(llm.tokenize) — vendor docs
llama.cpp CLI
Capabilities (4/4 rows usable):local.llamacpp.bench(eval.benchmark) — vendor docslocal.llamacpp.cli(provider.cli) — vendor docslocal.llamacpp.model_acquisition_cache(provider.models) — vendor docslocal.llamacpp.perplexity(eval.perplexity) — vendor docs
llama.cpp CLI Tools
Capabilities (15/20 rows usable):local.llamacpp.cli.batched(eval.benchmark) — vendor docslocal.llamacpp.cli.batched_bench(eval.benchmark) — vendor docslocal.llamacpp.cli.embedding(llm.embed) — vendor docslocal.llamacpp.cli.gguf_inspect(provider.diagnostics) — vendor docslocal.llamacpp.cli.gguf_split(file.split_merge) — vendor docslocal.llamacpp.cli.imatrix(provider.quantization) — vendor docslocal.llamacpp.cli.lookahead(llm.speculative_decoding) — vendor docslocal.llamacpp.cli.mtmd_cli(media.multimodal) — vendor docslocal.llamacpp.cli.parallel(eval.benchmark) — vendor docslocal.llamacpp.cli.passkey(eval.benchmark) — vendor docslocal.llamacpp.cli.quantize(provider.quantization) — vendor docslocal.llamacpp.cli.retrieval(llm.rag) — vendor docslocal.llamacpp.cli.simple(llm.complete) — vendor docslocal.llamacpp.cli.simple_chat(llm.chat) — vendor docslocal.llamacpp.cli.tokenize(llm.tokenize) — vendor docs
llama.cpp Evaluation
Capabilities (1/1 rows usable):local.llamacpp.eval.choice_logit_tasks(eval.benchmark) · model-dependent — vendor docs
llama.cpp Quant Types
Capabilities (7/7 rows usable):local.llamacpp.quant.bf16_f16_f32(provider.quantization) — vendor docslocal.llamacpp.quant.iq_extreme(provider.quantization) — vendor docslocal.llamacpp.quant.iq2(provider.quantization) — vendor docslocal.llamacpp.quant.iq3(provider.quantization) — vendor docslocal.llamacpp.quant.iq4(provider.quantization) — vendor docslocal.llamacpp.quant.k_quants(provider.quantization) — vendor docslocal.llamacpp.quant.legacy_q(provider.quantization) — vendor docs
llama.cpp Runtime
Capabilities (4/7 rows usable):local.llamacpp.runtime.cpu_memory(provider.runtime) — vendor docslocal.llamacpp.runtime.kv_cache_context(provider.context) — vendor docslocal.llamacpp.runtime.threads_batch(provider.runtime) — vendor docslocal.llamacpp.sampling_controls(llm.sampling) — vendor docs
llama.cpp Server Anthropic Format
Capabilities (2/2 rows usable):local.llamacpp.server.anthropic_count_tokens(anthropic.messages.count_tokens) — vendor docslocal.llamacpp.server.anthropic_messages(anthropic.messages) — vendor docs
llama.cpp Server Native
Capabilities (21/26 rows usable):local.llamacpp.server.apply_template(llm.chat_template) — vendor docslocal.llamacpp.server.auth_tls(provider.auth) — vendor docslocal.llamacpp.server.completion_native(llm.completion) — vendor docslocal.llamacpp.server.detokenize(llm.tokenize) — vendor docslocal.llamacpp.server.embeddings_native(llm.embedding) — vendor docslocal.llamacpp.server.gpu_backend(provider.gpu_offload) — vendor docslocal.llamacpp.server.grammar(llm.structured_output) — vendor docslocal.llamacpp.server.health(provider.health) — vendor docslocal.llamacpp.server.infill(llm.fim) · model-dependent — vendor docslocal.llamacpp.server.lora_adapters(tuning.lora_runtime) · model-dependent — vendor docslocal.llamacpp.server.metrics(provider.metrics) — vendor docslocal.llamacpp.server.parallel_batching(provider.parallel_decoding) — vendor docslocal.llamacpp.server.props(provider.runtime_config) — vendor docslocal.llamacpp.server.props_post(provider.runtime_config) — vendor docslocal.llamacpp.server.reasoning(llm.reasoning) · model-dependent — vendor docslocal.llamacpp.server.reranking(llm.rerank) — vendor docslocal.llamacpp.server.slots(provider.slots) — vendor docslocal.llamacpp.server.slots_save_restore(provider.slots) — vendor docslocal.llamacpp.server.speculative(llm.speculative_decoding) — vendor docslocal.llamacpp.server.tokenize(llm.tokenize) — vendor docslocal.llamacpp.server.webui(docs.webui) — vendor docs
llama.cpp Server OpenAI Format
Capabilities (2/2 rows usable):local.llamacpp.server.responses(llm.responses) — vendor docslocal.llamacpp.server.tools(tool.function_calling) · model-dependent — vendor docs
LocalAI Anthropic Format
Capabilities (1/1 rows usable):local.localai.anthropic_messages(anthropic.messages) — vendor docs
LocalAI Backends
Capabilities (1/23 rows usable):local.localai.backend.llamacpp(provider.backend) — vendor docs
LocalAI Galleries
Capabilities (2/6 rows usable):local.localai.gallery_available(provider.gallery) — vendor docslocal.localai.gallery_jobs(provider.gallery) — vendor docs
LocalAI Ollama Format
Capabilities (1/1 rows usable):local.localai.ollama_compat(llm.complete) — vendor docs
LocalAI OpenAI Format
Capabilities (3/16 rows usable):local.localai.openai_chat(llm.chat) — vendor docslocal.localai.openai_completions(llm.completion) — vendor docslocal.localai.openai_models(provider.models) — vendor docs
MLX Apple Platform
Capabilities (2/2 rows usable):local.mlx.platform.lazy_eval(provider.runtime) — vendor docslocal.mlx.platform.macos_unified_memory(provider.gpu_offload) — vendor docs
MLX Distributed
Capabilities (2/2 rows usable):local.mlx.distributed.launch(provider.distributed_inference) — vendor docslocal.mlx.distributed.primitives(provider.distributed_inference) — vendor docs
MLX Examples
Capabilities (2/12 rows usable):local.mlx.ex.bert(llm.embed) · model-dependent — vendor docslocal.mlx.ex.t5(llm.complete) · model-dependent — vendor docs
MLX Fast Kernels
Capabilities (5/5 rows usable):local.mlx.fast.metal_kernel(provider.runtime) — vendor docslocal.mlx.fast.quantized_matmul(provider.quantization) — vendor docslocal.mlx.fast.rms_norm(provider.runtime) — vendor docslocal.mlx.fast.rope(provider.runtime) — vendor docslocal.mlx.fast.scaled_dot_product_attention(provider.attention_backend) — vendor docs
MLX Optimizers
Capabilities (1/1 rows usable):local.mlx.optimizers(tuning.training_reference) — vendor docs
MLX Profiling
Capabilities (3/3 rows usable):local.mlx.metal.cache_limit(provider.gpu_offload) — vendor docslocal.mlx.metal.capture(provider.observability) — vendor docslocal.mlx.utils.tree(provider.runtime) — vendor docs
MLX-LM CLI
Capabilities (10/17 rows usable):local.mlx.lm.cli_benchmark(eval.benchmark) — vendor docslocal.mlx.lm.cli_cache_prompt(provider.slots) — vendor docslocal.mlx.lm.cli_chat(llm.chat) — vendor docslocal.mlx.lm.cli_convert(provider.models) — vendor docslocal.mlx.lm.cli_fuse(tuning.lora_runtime) — vendor docslocal.mlx.lm.cli_generate(llm.complete) — vendor docslocal.mlx.lm.cli_lora(tuning.lora_runtime) — vendor docslocal.mlx.lm.cli_manage(provider.models) — vendor docslocal.mlx.lm.cli_perplexity(eval.perplexity) · model-dependent — vendor docslocal.mlx.lm.cli_server(llm.chat) — vendor docs
MLX-LM Python SDK
Capabilities (7/9 rows usable):local.mlx.lm.kv_cache_quantized(provider.kv_cache) — vendor docslocal.mlx.lm.kv_cache_rotating(provider.kv_cache) — vendor docslocal.mlx.lm.py_generate(llm.complete) — vendor docslocal.mlx.lm.py_load(provider.models) — vendor docslocal.mlx.lm.py_prompt_cache(provider.slots) — vendor docslocal.mlx.lm.py_sample_utils(llm.sampling) — vendor docslocal.mlx.lm.py_stream_generate(llm.streaming) — vendor docs
MLX-LM Server
Capabilities (1/1 rows usable):local.mlx.lm.speculative(llm.speculative_decoding) — vendor docs
Ollama Anthropic Format
Capabilities (1/3 rows usable):local.ollama.anthropic_messages(anthropic.messages) — vendor docs
Ollama Blobs
Capabilities (2/2 rows usable):local.ollama.api_blobs_head(file.exists) — vendor docslocal.ollama.api_blobs_post(file.upload) — vendor docs
Ollama CLI
Capabilities (6/7 rows usable):local.ollama.cli_cp(provider.models) — vendor docslocal.ollama.cli_pull_rm_ls(provider.models) — vendor docslocal.ollama.cli_run(llm.chat) — vendor docslocal.ollama.cli_serve(provider.lifecycle) — vendor docslocal.ollama.cli_show(provider.admin.read) — vendor docslocal.ollama.cli_stop(provider.lifecycle) — vendor docs
Ollama Environment
Capabilities (13/15 rows usable):local.ollama.env.context_length(provider.context) — vendor docslocal.ollama.env.debug(provider.observability) — vendor docslocal.ollama.env.flash_attention(provider.attention_backend) — vendor docslocal.ollama.env.gpu_overhead(provider.gpu_offload) — vendor docslocal.ollama.env.host(provider.connectivity) — vendor docslocal.ollama.env.keep_alive(provider.lifecycle) — vendor docslocal.ollama.env.kv_cache_type(provider.kv_cache) — vendor docslocal.ollama.env.max_loaded(provider.lifecycle) — vendor docslocal.ollama.env.max_queue(provider.parallel_decoding) — vendor docslocal.ollama.env.models_dir(file.manage) — vendor docslocal.ollama.env.num_parallel(provider.parallel_decoding) — vendor docslocal.ollama.env.origins(provider.connectivity) — vendor docslocal.ollama.env.sched_spread(provider.gpu_offload) — vendor docs
Ollama Generate
Capabilities (6/6 rows usable):local.ollama.chat(llm.chat) — vendor docslocal.ollama.generate_context(llm.state.continue) — vendor docslocal.ollama.generate_keep_alive(provider.lifecycle) — vendor docslocal.ollama.generate_options_full(llm.sampling) — vendor docslocal.ollama.generate_raw(llm.complete) — vendor docslocal.ollama.generate_suffix(llm.fim) · model-dependent — vendor docs
Ollama Model Management
Capabilities (11/11 rows usable):local.ollama.api_create_adapters(tuning.lora_runtime) · model-dependent — vendor docslocal.ollama.api_create_quantize(provider.quantization) — vendor docslocal.ollama.api_create_safetensors(provider.models) · model-dependent — vendor docslocal.ollama.api_show(provider.admin.read) — vendor docslocal.ollama.copy(provider.models) — vendor docslocal.ollama.create(provider.models) — vendor docslocal.ollama.delete(provider.models) — vendor docslocal.ollama.ps(provider.lifecycle) — vendor docslocal.ollama.pull(provider.models) — vendor docslocal.ollama.push(provider.models) · model-dependent — vendor docslocal.ollama.tags(provider.models) — vendor docs
Ollama Modelfile
Capabilities (2/2 rows usable):local.ollama.context_length(provider.context) — vendor docslocal.ollama.modelfile(provider.modelfile) — vendor docs
Ollama OpenAI Format
Capabilities (6/6 rows usable):local.ollama.openai_chat(llm.chat) — vendor docslocal.ollama.openai_completions(llm.completions) — vendor docslocal.ollama.openai_embeddings(llm.embeddings) — vendor docslocal.ollama.openai_images(media.image_generation) · model-dependent — vendor docslocal.ollama.openai_models(provider.models) — vendor docslocal.ollama.openai_responses(llm.responses) — vendor docs
Ollama Server
Capabilities (1/1 rows usable):local.ollama.api_version(provider.health) — vendor docs
Ollama Streaming
Capabilities (1/1 rows usable):local.ollama.streaming(llm.streaming) — vendor docs
Ollama Structured Output
Concept guide: Structured Output — what this is, when to use it, and how to express it in workflow markdown.Capabilities (1/1 rows usable):
local.ollama.structured_outputs(llm.structured_output) — vendor docs
Ollama Thinking
Concept guide: Thinking & Reasoning — what this is, when to use it, and how to express it in workflow markdown.Capabilities (1/1 rows usable):
local.ollama.thinking(llm.thinking) · model-dependent — vendor docs
Ollama Vision
Capabilities (2/2 rows usable):local.ollama.generate_image_input(media.image_input) · model-dependent — vendor docslocal.ollama.vision(media.image_input) · model-dependent — vendor docs
ollama-js SDK
Capabilities (3/5 rows usable):local.ollama.js_abort_method(llm.cancel) — vendor docslocal.ollama.js_async_iterator(llm.streaming) — vendor docslocal.ollama.js_client_class(provider.admin.read) — vendor docs
ollama-python SDK
Capabilities (10/11 rows usable):local.ollama.python_async_client(provider.admin.read) — vendor docslocal.ollama.python_chat_method(llm.chat) — vendor docslocal.ollama.python_client_class(provider.admin.read) — vendor docslocal.ollama.python_copy_delete(provider.models) — vendor docslocal.ollama.python_create_modelfile(provider.models) — vendor docslocal.ollama.python_embed_method(llm.embed) — vendor docslocal.ollama.python_generate_method(llm.complete) — vendor docslocal.ollama.python_list_method(provider.models) — vendor docslocal.ollama.python_ps_method(provider.lifecycle) — vendor docslocal.ollama.python_show_method(provider.admin.read) — vendor docs
Retrieval/files/embeddings
Capabilities (4/4 rows usable):local.ollama.embed_dimensions(llm.embed) — vendor docslocal.ollama.embed_truncate(llm.embed) — vendor docslocal.ollama.embeddings(llm.embed) — vendor docslocal.ollama.embeddings_legacy(llm.embed) — vendor docs
Tools
Concept guide: Tools & MCP — what this is, when to use it, and how to express it in workflow markdown.Capabilities (1/1 rows usable):
local.ollama.tools(tool.call) · model-dependent — vendor docs