{"id":100,"date":"2026-09-23T11:24:43","date_gmt":"2026-09-23T11:24:43","guid":{"rendered":"https:\/\/docs.latw.ai\/uncategorized\/custom-openai-endpoint-2\/"},"modified":"2026-09-23T11:24:43","modified_gmt":"2026-09-23T11:24:43","slug":"custom-openai-endpoint-2","status":"publish","type":"post","link":"https:\/\/docs.latw.ai\/pl\/latw-for-polylang\/custom-openai-endpoint-2\/","title":{"rendered":"Using Your Own Model or Any OpenAI-Compatible API"},"content":{"rendered":"<p><em>Available in LATW AI Translation for Polylang <strong>PRO<\/strong>.<\/em><\/p>\n<p>Besides the built-in engines, the plugin can translate through <strong>any API that speaks OpenAI&#8217;s chat<br \/>\nformat<\/strong>. You give it an address, a model name and \u2014 if the service asks for one \u2014 a key.<\/p>\n<p>That one setting covers two very different situations:<\/p>\n<ul>\n<li><strong>A model you run yourself.<\/strong> On your own server, your own hardware, or a machine on your own<br \/>\nnetwork. Nothing leaves your infrastructure, and there is no per-word bill.<\/li>\n<li><strong>A gateway that resells other people&#8217;s models.<\/strong> OpenRouter, Groq and services like them expose<br \/>\nhundreds of models \u2014 open-weight and commercial alike \u2014 behind the same API, with one key and one<br \/>\ninvoice.<\/li>\n<\/ul>\n<p>Everything the other language-model engines do applies here: the translation prompt, the website<br \/>\ndescription and the glossary are all sent exactly as they are for OpenAI.<\/p>\n<h2>What you need before you start<\/h2>\n<ol>\n<li>An endpoint that answers at <code>POST \/chat\/completions<\/code> in OpenAI&#8217;s format. Every service and server<br \/>\nnamed on this page does.<\/li>\n<li>The exact model name that endpoint serves.<\/li>\n<li>A key, if the service requires one.<\/li>\n<li><strong>The address must be reachable from the web server that runs WordPress<\/strong> \u2014 not from your own<br \/>\ncomputer. A model running on your laptop is not reachable from a hosted WordPress site.<\/li>\n<\/ol>\n<h2>Setting it up<\/h2>\n<p>Go to <strong>AI Translation \u2192 Settings \u2192 General<\/strong> and pick <strong>Custom endpoint (OpenAI-compatible)<\/strong> in the<br \/>\n<strong>Engine<\/strong> list. The fields below it are the ones this engine uses.<\/p>\n<h3>Endpoint URL<\/h3>\n<p>The base URL of the API, <strong>without<\/strong> <code>\/chat\/completions<\/code> \u2014 that part is added for you:<\/p>\n<pre><code>https:\/\/openrouter.ai\/api\/v1\nhttps:\/\/api.groq.com\/openai\/v1\nhttp:\/\/127.0.0.1:11434\/v1<\/code><\/pre>\n<p>Most services print exactly this address in their own documentation, usually called the <em>base URL<\/em>.<br \/>\nIf you only have the full address ending in <code>\/chat\/completions<\/code>, you can paste that instead; it is<br \/>\nused as it is.<\/p>\n<h3>API Key<\/h3>\n<p>Sent as a Bearer token, the way OpenAI&#8217;s own API expects it. Leave it empty for a server that does<br \/>\nnot ask for one \u2014 a local Ollama or LM Studio, for instance.<\/p>\n<h3>Model<\/h3>\n<p>Typed in by hand, exactly as your endpoint names it. There is no dropdown, because the models are<br \/>\nwhatever your endpoint serves:<\/p>\n<pre><code>openai\/gpt-oss-120b            (OpenRouter)\nllama-3.3-70b-versatile        (Groq)\nllama3.1:8b                    (Ollama)<\/code><\/pre>\n<p>A name the endpoint does not recognise comes back as an error on the <strong>Logs<\/strong> page, usually saying<br \/>\nthe model does not exist.<\/p>\n<h3>Max output tokens<\/h3>\n<p>How long an answer the endpoint may return. The default of <strong>8192<\/strong> is enough for a typical page.<\/p>\n<p>This matters more than it looks. Several servers stop after a few hundred tokens unless they are<br \/>\nasked for more, and a translation that stops early arrives as an unfinished JSON object \u2014 which the<br \/>\nplugin reports as a failed translation rather than as a short one. If long posts fail while short<br \/>\nones succeed, raise this first.<\/p>\n<h3>JSON mode<\/h3>\n<p>Leave it on. It asks the endpoint to reply in JSON, which is the format the plugin reads answers in.<\/p>\n<p>Turn it off only if translations fail with an error mentioning <code>response_format<\/code> \u2014 a sign that this<br \/>\nparticular server does not support the setting. Translations still work without it: the prompt asks<br \/>\nfor JSON as well.<\/p>\n<h3>Reasoning effort<\/h3>\n<p>Off by default, and deliberately so. <strong>Reasoning Effort<\/strong>, the setting further down the same screen,<br \/>\nis only meaningful to models that reason before answering, and most servers have never heard of it. A<br \/>\nstrict server answers a setting it does not know with an error instead of ignoring it, which would<br \/>\nbreak every translation.<\/p>\n<p>Turn it on when you are pointing the plugin at a reasoning model \u2014 on OpenRouter or Groq, for<br \/>\nexample \u2014 and want the <strong>Reasoning Effort<\/strong> setting to reach it.<\/p>\n<h3>Price per 1M input tokens \/ Price per 1M output tokens<\/h3>\n<p>What your provider charges, in US dollars, copied from its pricing page. Per <strong>million<\/strong> tokens,<br \/>\nwhich is how these prices are normally published.<\/p>\n<p>The plugin does not know what your endpoint costs, so this is the only way for the <strong>History<\/strong> page<br \/>\nand the cost reports to show real numbers. Leave both at <strong>0<\/strong> for a model you host yourself, and the<br \/>\ntranslations are recorded as free \u2014 which they are.<\/p>\n<h2>Processing Mode<\/h2>\n<p>A custom endpoint is <strong>synchronous only<\/strong>. Batch and Background are services a vendor runs, and an<br \/>\nendpoint you configured yourself does not have them. If <strong>Processing<\/strong> is set to either, translations<br \/>\nrun synchronously anyway.<\/p>\n<p>In practice this means each translation is one HTTP request that the site waits for. A very long post<br \/>\non a slow model can therefore hit a PHP or web-server timeout and fail. If that happens, use a faster<br \/>\nmodel, or translate long posts in fewer, shorter passes.<\/p>\n<h2>Recipes<\/h2>\n<h3>OpenRouter<\/h3>\n<p>One key, hundreds of models, including free ones.<\/p>\n<table>\n<thead>\n<tr>\n<th>Field<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Endpoint URL<\/td>\n<td><code>https:\/\/openrouter.ai\/api\/v1<\/code><\/td>\n<\/tr>\n<tr>\n<td>API Key<\/td>\n<td>from <a href=\"https:\/\/openrouter.ai\/keys\">https:\/\/openrouter.ai\/keys<\/a><\/td>\n<\/tr>\n<tr>\n<td>Model<\/td>\n<td>the id shown on the model&#8217;s page, e.g. <code>openai\/gpt-oss-120b<\/code><\/td>\n<\/tr>\n<tr>\n<td>Prices<\/td>\n<td>from the same model page, per 1M tokens<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Model ids on OpenRouter always contain a slash \u2014 the vendor, then the model. Copy it exactly.<\/p>\n<h3>Groq<\/h3>\n<p>Known for being very fast, which suits synchronous translation well.<\/p>\n<table>\n<thead>\n<tr>\n<th>Field<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Endpoint URL<\/td>\n<td><code>https:\/\/api.groq.com\/openai\/v1<\/code><\/td>\n<\/tr>\n<tr>\n<td>API Key<\/td>\n<td>from <a href=\"https:\/\/console.groq.com\/keys\">https:\/\/console.groq.com\/keys<\/a><\/td>\n<\/tr>\n<tr>\n<td>Model<\/td>\n<td>e.g. <code>llama-3.3-70b-versatile<\/code><\/td>\n<\/tr>\n<tr>\n<td>Prices<\/td>\n<td>from Groq&#8217;s pricing page, per 1M tokens<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h3>Ollama on the same server as WordPress<\/h3>\n<table>\n<thead>\n<tr>\n<th>Field<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Endpoint URL<\/td>\n<td><code>http:\/\/127.0.0.1:11434\/v1<\/code><\/td>\n<\/tr>\n<tr>\n<td>API Key<\/td>\n<td>empty<\/td>\n<\/tr>\n<tr>\n<td>Model<\/td>\n<td>the name you pulled, e.g. <code>llama3.1:8b<\/code><\/td>\n<\/tr>\n<tr>\n<td>Prices<\/td>\n<td>0 and 0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Pull the model first (<code>ollama pull llama3.1:8b<\/code>) and make sure Ollama is running as a service, not<br \/>\njust in a terminal window. If WordPress runs inside a container, <code>127.0.0.1<\/code> points at the container<br \/>\nrather than at the host \u2014 use the address the container can reach.<\/p>\n<h3>vLLM, LM Studio, llama.cpp, LocalAI<\/h3>\n<p>All of them serve the same API. Use the address the server prints when it starts, with <code>\/v1<\/code> at the<br \/>\nend:<\/p>\n<pre><code>http:\/\/your-server:8000\/v1      (vLLM)\nhttp:\/\/127.0.0.1:1234\/v1        (LM Studio)<\/code><\/pre>\n<p>The model name is the one the server was started with. If you protected the server with a key, put it<br \/>\nin the <strong>API Key<\/strong> field; otherwise leave it empty.<\/p>\n<h2>Translation quality is now your responsibility<\/h2>\n<p>The built-in engines are models known to translate well. A custom endpoint is whatever you point it<br \/>\nat, and small local models vary enormously: some translate beautifully, others drop HTML tags, answer<br \/>\nin the wrong language, or return text that is not valid JSON at all.<\/p>\n<p>Before you turn it loose on a site:<\/p>\n<ul>\n<li>Translate <strong>one short post<\/strong> and read the result.<\/li>\n<li>Then translate <strong>one long post with formatting<\/strong> \u2014 headings, links, lists, a page builder if you<br \/>\nuse one \u2014 and check that the layout survived.<\/li>\n<li>Consider turning on the checks described in <em>Checking Translations Before They Are Saved<\/em>, which<br \/>\ncatch a mangled answer before it reaches your content.<\/li>\n<\/ul>\n<h2>When something goes wrong<\/h2>\n<p>Errors are recorded on the <strong>AI Translation \u2192 Logs<\/strong> page. Turning on <strong>Settings \u2192 Advanced \u2192 Debug<br \/>\nMode<\/strong> adds the full request and response, which is usually what identifies the problem.<\/p>\n<p>A gateway such as OpenRouter answers with an error of its own and keeps the real one \u2014 the message<br \/>\nfrom whoever actually ran the model \u2014 inside it. The plugin unpacks that, so the recorded error names<br \/>\nthe model, the reason and which company served it, and adds the setting to change where there is one.<\/p>\n<p>Unlike Background or Batch translations, a synchronous job that fails is not retried \u2014 you see the<br \/>\nprovider&#8217;s reason immediately, on the <strong>Translations<\/strong> and <strong>Logs<\/strong> screens, instead of waiting for<br \/>\nthe same request to be tried again.<\/p>\n<table>\n<thead>\n<tr>\n<th>What you see<\/th>\n<th>What it usually means<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Connection refused, or a cURL error<\/td>\n<td>The web server cannot reach the address. A local address only works if the model runs on the same machine as WordPress.<\/td>\n<\/tr>\n<tr>\n<td>404, or &#8220;not found&#8221;<\/td>\n<td>The URL is missing <code>\/v1<\/code>, or has a path the service does not use.<\/td>\n<\/tr>\n<tr>\n<td>401 or 403<\/td>\n<td>Key missing, wrong, or out of credit.<\/td>\n<\/tr>\n<tr>\n<td>An error about <code>response_format<\/code>, or &#8220;does not support structured outputs&#8221;<\/td>\n<td>The model cannot do JSON mode. Turn <strong>JSON mode<\/strong> off.<\/td>\n<\/tr>\n<tr>\n<td>An error mentioning <code>reasoning_effort<\/code><\/td>\n<td>The model is not a reasoning model. Turn <strong>Reasoning effort<\/strong> off.<\/td>\n<\/tr>\n<tr>\n<td>Translation was truncated<\/td>\n<td>Raise <strong>Max output tokens<\/strong>.<\/td>\n<\/tr>\n<tr>\n<td>&#8220;The translation engine did not return valid JSON&#8221;<\/td>\n<td>The model did not answer in JSON. Turn <strong>JSON mode<\/strong> on if it is off; if it is already on, the model is too weak for the job \u2014 try a larger one.<\/td>\n<\/tr>\n<tr>\n<td>An SSL or certificate error<\/td>\n<td>The endpoint needs a valid certificate. For a server on your own machine, use <code>http:\/\/<\/code> with a local address instead.<\/td>\n<\/tr>\n<tr>\n<td>Model does not exist<\/td>\n<td>The model name does not match what the endpoint serves.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Limits worth knowing<\/h2>\n<ul>\n<li><strong>Synchronous only.<\/strong> No Batch API, no Background mode, and therefore none of the 50% batch<br \/>\ndiscount.<\/li>\n<li><strong>Costs are what you type in.<\/strong> The plugin cannot read your provider&#8217;s prices, so the figures in<br \/>\nHistory are only as accurate as the two price fields.<\/li>\n<li><strong>One custom endpoint at a time<\/strong>, like every other engine.<\/li>\n<li><strong>No connection test.<\/strong> The first real translation is the test; run it on a short post.<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Available in LATW AI Translation for Polylang PRO. Besides the built-in engines, the plugin can translate through any API that speaks OpenAI&#8217;s chat format. You give\u2026<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_latwm_translated_slug":"","footnotes":""},"categories":[17],"tags":[],"class_list":["post-100","post","type-post","status-publish","format-standard","hentry","category-latw-for-polylang"],"_links":{"self":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/posts\/100","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/comments?post=100"}],"version-history":[{"count":0,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/posts\/100\/revisions"}],"wp:attachment":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/media?parent=100"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/categories?post=100"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/tags?post=100"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}