Using Your Own Model or Any OpenAI-Compatible API
On this page
- What you need before you start
- Setting it up
- Endpoint URL
- API Key
- Model
- Max output tokens
- JSON mode
- Reasoning effort
- Price per 1M input tokens / Price per 1M output tokens
- Processing Mode
- Recipes
- OpenRouter
- Groq
- Ollama on the same server as WordPress
- vLLM, LM Studio, llama.cpp, LocalAI
- Translation quality is now your responsibility
- When something goes wrong
- Limits worth knowing
- Developer reference
Available in PRO (version 2.4.0 and later).
Besides the built-in engines, the plugin can translate through any API that speaks OpenAI’s chat
format. You give it an address, a model name and – if the service asks for one – a key.
That one setting covers two very different situations:
- A model you run yourself. On your own server, your own hardware, or a machine on your own
network. Nothing leaves your infrastructure, and there is no per-word bill. - A gateway that resells other people’s models. OpenRouter, Groq and services like them expose
hundreds of models – open-weight and commercial alike – behind the same API, with one key and one
invoice.
Everything the other LLM engines do applies here: the translation prompt, the website description
and the glossary are all sent exactly as they are for OpenAI.
#What you need before you start
- An endpoint that answers at
POST /chat/completionsin OpenAI’s format. Every service and
server named on this page does. - The exact model name that endpoint serves.
- A key, if the service requires one.
- The address must be reachable from the web server that runs WordPress – not from your own
computer. A model running on your laptop is not reachable from a hosted WordPress site.
#Setting it up
Go to LATW Multilingual → Settings → AI and pick Custom endpoint (OpenAI-compatible) in the
Engine list. The fields below it are the ones this engine uses.
#Endpoint URL
The base URL of the API, without /chat/completions – that part is added for you:
https://openrouter.ai/api/v1
https://api.groq.com/openai/v1
http://127.0.0.1:11434/v1
Most services print exactly this address in their own documentation, usually called the base URL.
If you only have the full address ending in /chat/completions, you can paste that instead; it is
used as it is.
#API Key
Sent as a Bearer token, the way OpenAI’s own API expects it. Leave it empty for a server that does
not ask for one – a local Ollama or LM Studio, for instance.
#Model
Typed in by hand, exactly as your endpoint names it. There is no dropdown, because the models are
whatever your endpoint serves:
openai/gpt-oss-120b (OpenRouter)
llama-3.3-70b-versatile (Groq)
llama3.1:8b (Ollama)
A name the endpoint does not recognise comes back as an error on the History page, usually
saying the model does not exist.
The endpoint URL and the model are what the plugin needs before it will translate. Until both are
filled in, the translate buttons stay disabled and a notice at the top of the admin says the custom
endpoint is not configured.
#Max output tokens
How long an answer the endpoint may return. The default of 8192 is enough for a typical page.
This matters more than it looks. Several servers stop after a few hundred tokens unless they are
asked for more, and a translation that stops early would arrive as an unfinished answer. The plugin
reports it as truncated and names this setting, rather than saving half a page. If long posts fail
while short ones succeed, raise this first.
#JSON mode
Leave it on. It asks the endpoint to reply in JSON, which is the format the plugin reads answers in.
Turn it off only if translations fail with an error mentioning response_format or
structured outputs – a sign that this particular server or model does not support the setting.
Translations still work without it: the prompt asks for JSON as well.
#Reasoning effort
Off by default, and deliberately so. Reasoning Effort is only meaningful to models that reason
before answering, and most servers have never heard of it. A strict server answers a setting it
does not know with an error instead of ignoring it, which would break every translation.
Turn it on when you are pointing the plugin at a reasoning model – on OpenRouter or Groq, for
example. The Reasoning Effort setting then appears below and is sent with every request; while
this box is off it is hidden, because it would have no effect.
#Price per 1M input tokens / Price per 1M output tokens
What your provider charges, in US dollars, copied from its pricing page. Per million tokens,
which is how these prices are normally published.
The plugin does not know what your endpoint costs, so this is the only way for the History page
to show real numbers. Leave both at 0 for a model you host yourself, and the translations are
recorded as free – which they are.
Some servers do not report how many tokens a request used. Their translations are recorded with
0 tokens and a cost of 0, whatever the prices say.
#Processing Mode
A custom endpoint is synchronous only. Batch and Background are services a vendor runs, and an
endpoint you configured yourself does not have them. The Processing Mode list offers only
Synchronous while this engine is selected, and a mode saved earlier for another engine falls
back to Synchronous on its own.
In practice this means each translation is one HTTP request that the site waits for. A very long
post on a slow model can therefore hit a PHP or web-server timeout and fail. If that happens, use a
faster model, or translate long posts in fewer, shorter passes.
#Recipes
#OpenRouter
One key, hundreds of models, including free ones.
| Field | Value |
|---|---|
| Endpoint URL | https://openrouter.ai/api/v1 |
| API Key | from https://openrouter.ai/keys |
| Model | the id shown on the model’s page, e.g. openai/gpt-oss-120b |
| Prices | from the same model page, per 1M tokens |
Model ids on OpenRouter always contain a slash – the vendor, then the model. Copy it exactly.
#Groq
Known for being very fast, which suits synchronous translation well.
| Field | Value |
|---|---|
| Endpoint URL | https://api.groq.com/openai/v1 |
| API Key | from https://console.groq.com/keys |
| Model | e.g. llama-3.3-70b-versatile |
| Prices | from Groq’s pricing page, per 1M tokens |
#Ollama on the same server as WordPress
| Field | Value |
|---|---|
| Endpoint URL | http://127.0.0.1:11434/v1 |
| API Key | empty |
| Model | the name you pulled, e.g. llama3.1:8b |
| Prices | 0 and 0 |
Pull the model first (ollama pull llama3.1:8b) and make sure Ollama is running as a service, not
just in a terminal window. If WordPress runs inside a container, 127.0.0.1 points at the container
rather than at the host – use the address the container can reach.
#vLLM, LM Studio, llama.cpp, LocalAI
All of them serve the same API. Use the address the server prints when it starts, with /v1 at the
end:
http://your-server:8000/v1 (vLLM)
http://127.0.0.1:1234/v1 (LM Studio)
The model name is the one the server was started with. If you protected the server with a key, put
it in the API Key field; otherwise leave it empty.
#Translation quality is now your responsibility
The built-in engines are models known to translate well. A custom endpoint is whatever you point it
at, and small local models vary enormously: some translate beautifully, others drop HTML tags,
answer in the wrong language, or return text that is not valid JSON at all.
Before you turn it loose on a site:
- Translate one short post and read the result.
- Then translate one long post with formatting – headings, links, lists, an Elementor page if
you use one – and check that the layout survived.
#When something goes wrong
Failed translations show their error on the LATW Multilingual → History page, and the details
are written to LATW Multilingual → Logs. Turning on Settings → Advanced → Debug Mode records
more of each request, which is usually what identifies the problem.
A gateway such as OpenRouter answers with an error of its own and keeps the real one – the message
from whoever actually ran the model – inside it. The plugin unpacks that, so the recorded error
names the model, the reason and which company served it, and adds the setting to change where there
is one.
| What you see | What it usually means |
|---|---|
| Connection refused, or a cURL error | The web server cannot reach the address. A local address only works if the model runs on the same machine as WordPress. |
| 404, “not found”, or “Failed to parse the API response as JSON” | The URL is missing /v1, or has a path the service does not use. |
| 401 or 403 | Key missing, wrong, or out of credit. |
An error about response_format, or “does not support structured outputs” |
The model cannot do JSON mode. Turn JSON mode off. |
An error mentioning reasoning_effort |
The model is not a reasoning model. Turn Reasoning effort off. |
| The translation was truncated | Raise Max output tokens. |
| “Failed to parse the model response as JSON” | The model did not answer in JSON. Turn JSON mode on if it is off; if it is already on, the model is too weak for the job – try a larger one. |
| An SSL or certificate error | The endpoint needs a valid certificate. For a server on your own machine, use http:// with a local address instead. |
| Model does not exist | The model name does not match what the endpoint serves. |
#Limits worth knowing
- Synchronous only. No Batch API, no Background mode, and therefore none of the 50% batch
discount. - Costs are what you type in. The plugin cannot read your provider’s prices, so the figures in
History are only as accurate as the two price fields. - One custom endpoint at a time, like every other engine.
- No connection test. The first real translation is the test; run it on a short post.
- Long model names are shortened in History. The full name is always sent to your endpoint, but
the History page shows at most the first 60 characters of it.
#Developer reference
The request body can be changed before it is sent – to add temperature, top_p or a vendor
extension the settings screen does not offer:
add_filter( 'latwmp_custom_openai_request_body', function ( $body, $model, $prompt ) {
$body['temperature'] = 0.2;
return $body;
}, 10, 3 );
Whatever the filter returns is sent as-is. The engine’s provider id is custom_openai; its settings
live in the shared latwm_ai_settings option under the custom_openai_* keys, and are removed when
the PRO add-on is uninstalled.