Skip to content

Using Your Own Model or Any OpenAI-Compatible API

On this page

Available in LATW AI Translator for WPML PRO (version 3.0.1 and later).

Besides the built-in engines, the plugin can translate through any API that speaks OpenAI’s chat
format
. You give it an address, a model name and – if the service asks for one – a key.

That one setting covers two very different situations:

  • A model you run yourself. On your own server, your own hardware, or a machine on your own
    network. Nothing leaves your infrastructure, and there is no per-word bill.
  • A gateway that resells other people’s models. OpenRouter, Groq and services like them expose
    hundreds of models – open-weight and commercial alike – behind the same API, with one key and one
    invoice.

Everything the other LLM engines do applies here: the translation prompt, the website description
and the glossary are all sent exactly as they are for OpenAI.

#What you need before you start

  1. An endpoint that answers at POST /chat/completions in OpenAI’s format. Every service and
    server named on this page does.
  2. The exact model name that endpoint serves.
  3. A key, if the service requires one.
  4. The address must be reachable from the web server that runs WordPress – not from your own
    computer. A model running on your laptop is not reachable from a hosted WordPress site.

#Setting it up

Go to AI Translator → Settings → General and pick Custom endpoint (OpenAI-compatible) in the
Engine list. The fields below it are the ones this engine uses.

#Endpoint URL

The base URL of the API, without /chat/completions – that part is added for you:

https://openrouter.ai/api/v1
https://api.groq.com/openai/v1
http://127.0.0.1:11434/v1

Most services print exactly this address in their own documentation, usually called the base URL.
If you only have the full address ending in /chat/completions, you can paste that instead; it is
used as it is.

#API Key

Sent as a Bearer token, the way OpenAI’s own API expects it. Leave it empty for a server that does
not ask for one – a local Ollama or LM Studio, for instance.

#Model

Typed in by hand, exactly as your endpoint names it. There is no dropdown, because the models are
whatever your endpoint serves:

openai/gpt-oss-120b            (OpenRouter)
llama-3.3-70b-versatile        (Groq)
llama3.1:8b                    (Ollama)

A name the endpoint does not recognise comes back as an error on the Logs page, usually saying
the model does not exist.

#Max output tokens

How long an answer the endpoint may return. The default of 8192 is enough for a typical page.

This matters more than it looks. Several servers stop after a few hundred tokens unless they are
asked for more, and a translation that stops early arrives as an unfinished JSON object – which the
plugin reports as a failed translation rather than as a short one. If long posts fail while short
ones succeed, raise this first.

#JSON mode

Leave it on. It asks the endpoint to reply in JSON, which is the format the plugin reads answers in.

Turn it off only if translations fail with an error mentioning response_format – a sign that this
particular server does not support the setting. Translations still work without it: the prompt asks
for JSON as well.

#Reasoning effort

Off by default, and deliberately so. Reasoning Effort, the setting further down the same screen,
is only meaningful to models that reason before answering, and most servers have never heard of it.
A strict server answers a setting it does not know with an error instead of ignoring it, which would
break every translation.

Turn it on when you are pointing the plugin at a reasoning model – on OpenRouter or Groq, for
example – and want the Reasoning Effort setting to reach it.

#Price per 1M input tokens / Price per 1M output tokens

What your provider charges, in US dollars, copied from its pricing page. Per million tokens,
which is how these prices are normally published.

The plugin does not know what your endpoint costs, so this is the only way for the History page
and the cost reports to show real numbers. Leave both at 0 for a model you host yourself, and
the translations are recorded as free – which they are.

#Processing Mode

A custom endpoint is synchronous only. Batch and Background are services a vendor runs, and an
endpoint you configured yourself does not have them. If the Processing Mode is set to either,
translations run synchronously anyway.

In practice this means each translation is one HTTP request that the site waits for. A very long
post on a slow model can therefore hit a PHP or web-server timeout and be retried from scratch. If
that happens, use a faster model, or translate long posts in fewer, shorter passes.

#Recipes

#OpenRouter

One key, hundreds of models, including free ones.

Field Value
Endpoint URL https://openrouter.ai/api/v1
API Key from https://openrouter.ai/keys
Model the id shown on the model’s page, e.g. openai/gpt-oss-120b
Prices from the same model page, per 1M tokens

Model ids on OpenRouter always contain a slash – the vendor, then the model. Copy it exactly.

#Groq

Known for being very fast, which suits synchronous translation well.

Field Value
Endpoint URL https://api.groq.com/openai/v1
API Key from https://console.groq.com/keys
Model e.g. llama-3.3-70b-versatile
Prices from Groq’s pricing page, per 1M tokens

#Ollama on the same server as WordPress

Field Value
Endpoint URL http://127.0.0.1:11434/v1
API Key empty
Model the name you pulled, e.g. llama3.1:8b
Prices 0 and 0

Pull the model first (ollama pull llama3.1:8b) and make sure Ollama is running as a service, not
just in a terminal window. If WordPress runs inside a container, 127.0.0.1 points at the container
rather than at the host – use the address the container can reach.

#vLLM, LM Studio, llama.cpp, LocalAI

All of them serve the same API. Use the address the server prints when it starts, with /v1 at the
end:

http://your-server:8000/v1      (vLLM)
http://127.0.0.1:1234/v1        (LM Studio)

The model name is the one the server was started with. If you protected the server with a key, put
it in the API Key field; otherwise leave it empty.

#Translation quality is now your responsibility

The built-in engines are models known to translate well. A custom endpoint is whatever you point it
at, and small local models vary enormously: some translate beautifully, others drop HTML tags,
answer in the wrong language, or return text that is not valid JSON at all.

Before you turn it loose on a site:

  • Translate one short post and read the result.
  • Then translate one long post with formatting – headings, links, lists, a page builder if you
    use one – and check that the layout survived.
  • Consider turning on the checks described in Checking Translations Before They Are Saved, which
    catch a mangled answer before it reaches your content.

#When something goes wrong

Errors are recorded on the AI Translator → Logs page. Turning on Settings → Advanced → Debug
Mode
adds the full request and response, which is usually what identifies the problem.

A gateway such as OpenRouter answers with an error of its own and keeps the real one – the message
from whoever actually ran the model – inside it. The plugin unpacks that, so the recorded error
names the model, the reason and which company served it, and adds the setting to change where there
is one.

A request the provider rejects – an unknown model, a setting it does not accept, a key it does not
like – is not retried. You see the reason on the first attempt, on the Translations and
History screens, instead of waiting for three identical attempts to fail. Timeouts, rate limits
and provider outages are still retried, because those do pass.

What you see What it usually means
Connection refused, or a cURL error The web server cannot reach the address. A local address only works if the model runs on the same machine as WordPress.
404, or “not found” The URL is missing /v1, or has a path the service does not use.
401 or 403 Key missing, wrong, or out of credit.
An error about response_format, or “does not support structured outputs” The model cannot do JSON mode. Turn JSON mode off.
An error mentioning reasoning_effort The model is not a reasoning model. Turn Reasoning effort off.
Translation was truncated Raise Max output tokens.
“Failed to parse JSON response” The model did not answer in JSON. Turn JSON mode on if it is off; if it is already on, the model is too weak for the job – try a larger one.
An SSL or certificate error The endpoint needs a valid certificate. For a server on your own machine, use http:// with a local address instead.
Model does not exist The model name does not match what the endpoint serves.

#Limits worth knowing

  • Synchronous only. No Batch API, no Background mode, and therefore none of the 50% batch
    discount.
  • Costs are what you type in. The plugin cannot read your provider’s prices, so the figures in
    History are only as accurate as the two price fields.
  • One custom endpoint at a time, like every other engine.
  • No connection test. The first real translation is the test; run it on a short post.