Our take
DekaLLM is a small, budget-oriented inference host with a tight catalogue of seven tracked offers. It stands out for very low input pricing on Mistral Nemo and high throughput on one mid-size model, though its compliance documentation is thin and speeds vary wildly across the catalogue.
Use this for low-cost inference on Mistral Nemo, or for high-throughput work on Qwen3 6 where the output price is acceptable. Skip it if you need verified SOC 2, an EU endpoint, or a zero-retention guarantee; if you need a deep catalogue to choose from; or if your workload requires consistently fast throughput across models.
- Very low input price on Mistral Nemo — the cheapest input in its own catalogue.
- Highest throughput in its own catalogue on Qwen3 6 35B A3B at 163 tokens per second, roughly double its next-fastest offer.
- Does not train on prompts — confirmed in its privacy policy.
- Extremely slow throughput on Nemotron 3 Super at 7 tokens per second, the slowest in its catalogue by a wide margin.
- Steep output pricing on two offers and thin compliance documentation: SOC 2, EU endpoint, headquarters country and zero-retention are all unverified in our data.
- Small catalogue depth — only seven tracked offers.
Point your tools here
We have not recorded what a router needs for DekaLLM yet — no base URL and no docs link on file. That is our gap, not a sign DekaLLM has no API; its own documentation is the place to look until we close it.
Privacy & data handling
These answers cover requests sent to DekaLLM through OpenRouter, as OpenRouter records them, last read 58 min ago. For DekaLLM's own API we hold no answer: check its terms. SOC 2 and the data processing agreement below are hand-checked, not part of that sync.
Prompt retention: none.