Skip to content

Using Your Own Model or Any OpenAI-Compatible API

On this page

Available in LATW AI Translation for Polylang PRO.

Besides the built-in engines, the plugin can translate through any API that speaks OpenAI’s chat
format
. You give it an address, a model name and — if the service asks for one — a key.

That one setting covers two very different situations:

  • A model you run yourself. On your own server, your own hardware, or a machine on your own
    network. Nothing leaves your infrastructure, and there is no per-word bill.
  • A gateway that resells other people’s models. OpenRouter, Groq and services like them expose
    hundreds of models — open-weight and commercial alike — behind the same API, with one key and one
    invoice.

Everything the other language-model engines do applies here: the translation prompt, the website
description and the glossary are all sent exactly as they are for OpenAI.

#What you need before you start

  1. An endpoint that answers at POST /chat/completions in OpenAI’s format. Every service and server
    named on this page does.
  2. The exact model name that endpoint serves.
  3. A key, if the service requires one.
  4. The address must be reachable from the web server that runs WordPress — not from your own
    computer. A model running on your laptop is not reachable from a hosted WordPress site.

#Setting it up

Go to AI Translation → Settings → General and pick Custom endpoint (OpenAI-compatible) in the
Engine list. The fields below it are the ones this engine uses.

#Endpoint URL

The base URL of the API, without /chat/completions — that part is added for you:

https://openrouter.ai/api/v1
https://api.groq.com/openai/v1
http://127.0.0.1:11434/v1

Most services print exactly this address in their own documentation, usually called the base URL.
If you only have the full address ending in /chat/completions, you can paste that instead; it is
used as it is.

#API Key

Sent as a Bearer token, the way OpenAI’s own API expects it. Leave it empty for a server that does
not ask for one — a local Ollama or LM Studio, for instance.

#Model

Typed in by hand, exactly as your endpoint names it. There is no dropdown, because the models are
whatever your endpoint serves:

openai/gpt-oss-120b            (OpenRouter)
llama-3.3-70b-versatile        (Groq)
llama3.1:8b                    (Ollama)

A name the endpoint does not recognise comes back as an error on the Logs page, usually saying
the model does not exist.

#Max output tokens

How long an answer the endpoint may return. The default of 8192 is enough for a typical page.

This matters more than it looks. Several servers stop after a few hundred tokens unless they are
asked for more, and a translation that stops early arrives as an unfinished JSON object — which the
plugin reports as a failed translation rather than as a short one. If long posts fail while short
ones succeed, raise this first.

#JSON mode

Leave it on. It asks the endpoint to reply in JSON, which is the format the plugin reads answers in.

Turn it off only if translations fail with an error mentioning response_format — a sign that this
particular server does not support the setting. Translations still work without it: the prompt asks
for JSON as well.

#Reasoning effort

Off by default, and deliberately so. Reasoning Effort, the setting further down the same screen,
is only meaningful to models that reason before answering, and most servers have never heard of it. A
strict server answers a setting it does not know with an error instead of ignoring it, which would
break every translation.

Turn it on when you are pointing the plugin at a reasoning model — on OpenRouter or Groq, for
example — and want the Reasoning Effort setting to reach it.

#Price per 1M input tokens / Price per 1M output tokens

What your provider charges, in US dollars, copied from its pricing page. Per million tokens,
which is how these prices are normally published.

The plugin does not know what your endpoint costs, so this is the only way for the History page
and the cost reports to show real numbers. Leave both at 0 for a model you host yourself, and the
translations are recorded as free — which they are.

#Processing Mode

A custom endpoint is synchronous only. Batch and Background are services a vendor runs, and an
endpoint you configured yourself does not have them. If Processing is set to either, translations
run synchronously anyway.

In practice this means each translation is one HTTP request that the site waits for. A very long post
on a slow model can therefore hit a PHP or web-server timeout and fail. If that happens, use a faster
model, or translate long posts in fewer, shorter passes.

#Recipes

#OpenRouter

One key, hundreds of models, including free ones.

Field Value
Endpoint URL https://openrouter.ai/api/v1
API Key from https://openrouter.ai/keys
Model the id shown on the model’s page, e.g. openai/gpt-oss-120b
Prices from the same model page, per 1M tokens

Model ids on OpenRouter always contain a slash — the vendor, then the model. Copy it exactly.

#Groq

Known for being very fast, which suits synchronous translation well.

Field Value
Endpoint URL https://api.groq.com/openai/v1
API Key from https://console.groq.com/keys
Model e.g. llama-3.3-70b-versatile
Prices from Groq’s pricing page, per 1M tokens

#Ollama on the same server as WordPress

Field Value
Endpoint URL http://127.0.0.1:11434/v1
API Key empty
Model the name you pulled, e.g. llama3.1:8b
Prices 0 and 0

Pull the model first (ollama pull llama3.1:8b) and make sure Ollama is running as a service, not
just in a terminal window. If WordPress runs inside a container, 127.0.0.1 points at the container
rather than at the host — use the address the container can reach.

#vLLM, LM Studio, llama.cpp, LocalAI

All of them serve the same API. Use the address the server prints when it starts, with /v1 at the
end:

http://your-server:8000/v1      (vLLM)
http://127.0.0.1:1234/v1        (LM Studio)

The model name is the one the server was started with. If you protected the server with a key, put it
in the API Key field; otherwise leave it empty.

#Translation quality is now your responsibility

The built-in engines are models known to translate well. A custom endpoint is whatever you point it
at, and small local models vary enormously: some translate beautifully, others drop HTML tags, answer
in the wrong language, or return text that is not valid JSON at all.

Before you turn it loose on a site:

  • Translate one short post and read the result.
  • Then translate one long post with formatting — headings, links, lists, a page builder if you
    use one — and check that the layout survived.
  • Consider turning on the checks described in Checking Translations Before They Are Saved, which
    catch a mangled answer before it reaches your content.

#When something goes wrong

Errors are recorded on the AI Translation → Logs page. Turning on Settings → Advanced → Debug
Mode
adds the full request and response, which is usually what identifies the problem.

A gateway such as OpenRouter answers with an error of its own and keeps the real one — the message
from whoever actually ran the model — inside it. The plugin unpacks that, so the recorded error names
the model, the reason and which company served it, and adds the setting to change where there is one.

Unlike Background or Batch translations, a synchronous job that fails is not retried — you see the
provider’s reason immediately, on the Translations and Logs screens, instead of waiting for
the same request to be tried again.

What you see What it usually means
Connection refused, or a cURL error The web server cannot reach the address. A local address only works if the model runs on the same machine as WordPress.
404, or “not found” The URL is missing /v1, or has a path the service does not use.
401 or 403 Key missing, wrong, or out of credit.
An error about response_format, or “does not support structured outputs” The model cannot do JSON mode. Turn JSON mode off.
An error mentioning reasoning_effort The model is not a reasoning model. Turn Reasoning effort off.
Translation was truncated Raise Max output tokens.
“The translation engine did not return valid JSON” The model did not answer in JSON. Turn JSON mode on if it is off; if it is already on, the model is too weak for the job — try a larger one.
An SSL or certificate error The endpoint needs a valid certificate. For a server on your own machine, use http:// with a local address instead.
Model does not exist The model name does not match what the endpoint serves.

#Limits worth knowing

  • Synchronous only. No Batch API, no Background mode, and therefore none of the 50% batch
    discount.
  • Costs are what you type in. The plugin cannot read your provider’s prices, so the figures in
    History are only as accurate as the two price fields.
  • One custom endpoint at a time, like every other engine.
  • No connection test. The first real translation is the test; run it on a short post.