{"id":110,"date":"2026-09-24T09:57:26","date_gmt":"2026-09-24T09:57:26","guid":{"rendered":"https:\/\/docs.latw.ai\/uncategorized\/custom-openai-endpoint-3\/"},"modified":"2026-09-24T09:57:26","modified_gmt":"2026-09-24T09:57:26","slug":"custom-openai-endpoint-3","status":"publish","type":"post","link":"https:\/\/docs.latw.ai\/pl\/latw-multilingual\/custom-openai-endpoint-3\/","title":{"rendered":"Using Your Own Model or Any OpenAI-Compatible API"},"content":{"rendered":"<p><em>Available in <strong>PRO<\/strong> (version 2.4.0 and later).<\/em><\/p>\n<p>Besides the built-in engines, the plugin can translate through <strong>any API that speaks OpenAI&#8217;s chat<br \/>\nformat<\/strong>. You give it an address, a model name and &#8211; if the service asks for one &#8211; a key.<\/p>\n<p>That one setting covers two very different situations:<\/p>\n<ul>\n<li><strong>A model you run yourself.<\/strong> On your own server, your own hardware, or a machine on your own<br \/>\nnetwork. Nothing leaves your infrastructure, and there is no per-word bill.<\/li>\n<li><strong>A gateway that resells other people&#8217;s models.<\/strong> OpenRouter, Groq and services like them expose<br \/>\nhundreds of models &#8211; open-weight and commercial alike &#8211; behind the same API, with one key and one<br \/>\ninvoice.<\/li>\n<\/ul>\n<p>Everything the other LLM engines do applies here: the translation prompt, the website description<br \/>\nand the glossary are all sent exactly as they are for OpenAI.<\/p>\n<h2>What you need before you start<\/h2>\n<ol>\n<li>An endpoint that answers at <code>POST \/chat\/completions<\/code> in OpenAI&#8217;s format. Every service and<br \/>\nserver named on this page does.<\/li>\n<li>The exact model name that endpoint serves.<\/li>\n<li>A key, if the service requires one.<\/li>\n<li><strong>The address must be reachable from the web server that runs WordPress<\/strong> &#8211; not from your own<br \/>\ncomputer. A model running on your laptop is not reachable from a hosted WordPress site.<\/li>\n<\/ol>\n<h2>Setting it up<\/h2>\n<p>Go to <strong>LATW Multilingual \u2192 Settings \u2192 AI<\/strong> and pick <strong>Custom endpoint (OpenAI-compatible)<\/strong> in the<br \/>\n<strong>Engine<\/strong> list. The fields below it are the ones this engine uses.<\/p>\n<h3>Endpoint URL<\/h3>\n<p>The base URL of the API, <strong>without<\/strong> <code>\/chat\/completions<\/code> &#8211; that part is added for you:<\/p>\n<pre><code>https:\/\/openrouter.ai\/api\/v1\nhttps:\/\/api.groq.com\/openai\/v1\nhttp:\/\/127.0.0.1:11434\/v1<\/code><\/pre>\n<p>Most services print exactly this address in their own documentation, usually called the <em>base URL<\/em>.<br \/>\nIf you only have the full address ending in <code>\/chat\/completions<\/code>, you can paste that instead; it is<br \/>\nused as it is.<\/p>\n<h3>API Key<\/h3>\n<p>Sent as a Bearer token, the way OpenAI&#8217;s own API expects it. Leave it empty for a server that does<br \/>\nnot ask for one &#8211; a local Ollama or LM Studio, for instance.<\/p>\n<h3>Model<\/h3>\n<p>Typed in by hand, exactly as your endpoint names it. There is no dropdown, because the models are<br \/>\nwhatever your endpoint serves:<\/p>\n<pre><code>openai\/gpt-oss-120b            (OpenRouter)\nllama-3.3-70b-versatile        (Groq)\nllama3.1:8b                    (Ollama)<\/code><\/pre>\n<p>A name the endpoint does not recognise comes back as an error on the <strong>History<\/strong> page, usually<br \/>\nsaying the model does not exist.<\/p>\n<p>The endpoint URL and the model are what the plugin needs before it will translate. Until both are<br \/>\nfilled in, the translate buttons stay disabled and a notice at the top of the admin says the custom<br \/>\nendpoint is not configured.<\/p>\n<h3>Max output tokens<\/h3>\n<p>How long an answer the endpoint may return. The default of <strong>8192<\/strong> is enough for a typical page.<\/p>\n<p>This matters more than it looks. Several servers stop after a few hundred tokens unless they are<br \/>\nasked for more, and a translation that stops early would arrive as an unfinished answer. The plugin<br \/>\nreports it as <em>truncated<\/em> and names this setting, rather than saving half a page. If long posts fail<br \/>\nwhile short ones succeed, raise this first.<\/p>\n<h3>JSON mode<\/h3>\n<p>Leave it on. It asks the endpoint to reply in JSON, which is the format the plugin reads answers in.<\/p>\n<p>Turn it off only if translations fail with an error mentioning <code>response_format<\/code> or<br \/>\n<em>structured outputs<\/em> &#8211; a sign that this particular server or model does not support the setting.<br \/>\nTranslations still work without it: the prompt asks for JSON as well.<\/p>\n<h3>Reasoning effort<\/h3>\n<p>Off by default, and deliberately so. <strong>Reasoning Effort<\/strong> is only meaningful to models that reason<br \/>\nbefore answering, and most servers have never heard of it. A strict server answers a setting it<br \/>\ndoes not know with an error instead of ignoring it, which would break every translation.<\/p>\n<p>Turn it on when you are pointing the plugin at a reasoning model &#8211; on OpenRouter or Groq, for<br \/>\nexample. The <strong>Reasoning Effort<\/strong> setting then appears below and is sent with every request; while<br \/>\nthis box is off it is hidden, because it would have no effect.<\/p>\n<h3>Price per 1M input tokens \/ Price per 1M output tokens<\/h3>\n<p>What your provider charges, in US dollars, copied from its pricing page. Per <strong>million<\/strong> tokens,<br \/>\nwhich is how these prices are normally published.<\/p>\n<p>The plugin does not know what your endpoint costs, so this is the only way for the <strong>History<\/strong> page<br \/>\nto show real numbers. Leave both at <strong>0<\/strong> for a model you host yourself, and the translations are<br \/>\nrecorded as free &#8211; which they are.<\/p>\n<p>Some servers do not report how many tokens a request used. Their translations are recorded with<br \/>\n0 tokens and a cost of 0, whatever the prices say.<\/p>\n<h2>Processing Mode<\/h2>\n<p>A custom endpoint is <strong>synchronous only<\/strong>. Batch and Background are services a vendor runs, and an<br \/>\nendpoint you configured yourself does not have them. The <strong>Processing Mode<\/strong> list offers only<br \/>\n<strong>Synchronous<\/strong> while this engine is selected, and a mode saved earlier for another engine falls<br \/>\nback to Synchronous on its own.<\/p>\n<p>In practice this means each translation is one HTTP request that the site waits for. A very long<br \/>\npost on a slow model can therefore hit a PHP or web-server timeout and fail. If that happens, use a<br \/>\nfaster model, or translate long posts in fewer, shorter passes.<\/p>\n<h2>Recipes<\/h2>\n<h3>OpenRouter<\/h3>\n<p>One key, hundreds of models, including free ones.<\/p>\n<table>\n<thead>\n<tr>\n<th>Field<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Endpoint URL<\/td>\n<td><code>https:\/\/openrouter.ai\/api\/v1<\/code><\/td>\n<\/tr>\n<tr>\n<td>API Key<\/td>\n<td>from <a href=\"https:\/\/openrouter.ai\/keys\">https:\/\/openrouter.ai\/keys<\/a><\/td>\n<\/tr>\n<tr>\n<td>Model<\/td>\n<td>the id shown on the model&#8217;s page, e.g. <code>openai\/gpt-oss-120b<\/code><\/td>\n<\/tr>\n<tr>\n<td>Prices<\/td>\n<td>from the same model page, per 1M tokens<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Model ids on OpenRouter always contain a slash &#8211; the vendor, then the model. Copy it exactly.<\/p>\n<h3>Groq<\/h3>\n<p>Known for being very fast, which suits synchronous translation well.<\/p>\n<table>\n<thead>\n<tr>\n<th>Field<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Endpoint URL<\/td>\n<td><code>https:\/\/api.groq.com\/openai\/v1<\/code><\/td>\n<\/tr>\n<tr>\n<td>API Key<\/td>\n<td>from <a href=\"https:\/\/console.groq.com\/keys\">https:\/\/console.groq.com\/keys<\/a><\/td>\n<\/tr>\n<tr>\n<td>Model<\/td>\n<td>e.g. <code>llama-3.3-70b-versatile<\/code><\/td>\n<\/tr>\n<tr>\n<td>Prices<\/td>\n<td>from Groq&#8217;s pricing page, per 1M tokens<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h3>Ollama on the same server as WordPress<\/h3>\n<table>\n<thead>\n<tr>\n<th>Field<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Endpoint URL<\/td>\n<td><code>http:\/\/127.0.0.1:11434\/v1<\/code><\/td>\n<\/tr>\n<tr>\n<td>API Key<\/td>\n<td>empty<\/td>\n<\/tr>\n<tr>\n<td>Model<\/td>\n<td>the name you pulled, e.g. <code>llama3.1:8b<\/code><\/td>\n<\/tr>\n<tr>\n<td>Prices<\/td>\n<td>0 and 0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Pull the model first (<code>ollama pull llama3.1:8b<\/code>) and make sure Ollama is running as a service, not<br \/>\njust in a terminal window. If WordPress runs inside a container, <code>127.0.0.1<\/code> points at the container<br \/>\nrather than at the host &#8211; use the address the container can reach.<\/p>\n<h3>vLLM, LM Studio, llama.cpp, LocalAI<\/h3>\n<p>All of them serve the same API. Use the address the server prints when it starts, with <code>\/v1<\/code> at the<br \/>\nend:<\/p>\n<pre><code>http:\/\/your-server:8000\/v1      (vLLM)\nhttp:\/\/127.0.0.1:1234\/v1        (LM Studio)<\/code><\/pre>\n<p>The model name is the one the server was started with. If you protected the server with a key, put<br \/>\nit in the <strong>API Key<\/strong> field; otherwise leave it empty.<\/p>\n<h2>Translation quality is now your responsibility<\/h2>\n<p>The built-in engines are models known to translate well. A custom endpoint is whatever you point it<br \/>\nat, and small local models vary enormously: some translate beautifully, others drop HTML tags,<br \/>\nanswer in the wrong language, or return text that is not valid JSON at all.<\/p>\n<p>Before you turn it loose on a site:<\/p>\n<ul>\n<li>Translate <strong>one short post<\/strong> and read the result.<\/li>\n<li>Then translate <strong>one long post with formatting<\/strong> &#8211; headings, links, lists, an Elementor page if<br \/>\nyou use one &#8211; and check that the layout survived.<\/li>\n<\/ul>\n<h2>When something goes wrong<\/h2>\n<p>Failed translations show their error on the <strong>LATW Multilingual \u2192 History<\/strong> page, and the details<br \/>\nare written to <strong>LATW Multilingual \u2192 Logs<\/strong>. Turning on <strong>Settings \u2192 Advanced \u2192 Debug Mode<\/strong> records<br \/>\nmore of each request, which is usually what identifies the problem.<\/p>\n<p>A gateway such as OpenRouter answers with an error of its own and keeps the real one &#8211; the message<br \/>\nfrom whoever actually ran the model &#8211; inside it. The plugin unpacks that, so the recorded error<br \/>\nnames the model, the reason and which company served it, and adds the setting to change where there<br \/>\nis one.<\/p>\n<table>\n<thead>\n<tr>\n<th>What you see<\/th>\n<th>What it usually means<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Connection refused, or a cURL error<\/td>\n<td>The web server cannot reach the address. A local address only works if the model runs on the same machine as WordPress.<\/td>\n<\/tr>\n<tr>\n<td>404, &#8220;not found&#8221;, or &#8220;Failed to parse the API response as JSON&#8221;<\/td>\n<td>The URL is missing <code>\/v1<\/code>, or has a path the service does not use.<\/td>\n<\/tr>\n<tr>\n<td>401 or 403<\/td>\n<td>Key missing, wrong, or out of credit.<\/td>\n<\/tr>\n<tr>\n<td>An error about <code>response_format<\/code>, or &#8220;does not support structured outputs&#8221;<\/td>\n<td>The model cannot do JSON mode. Turn <strong>JSON mode<\/strong> off.<\/td>\n<\/tr>\n<tr>\n<td>An error mentioning <code>reasoning_effort<\/code><\/td>\n<td>The model is not a reasoning model. Turn <strong>Reasoning effort<\/strong> off.<\/td>\n<\/tr>\n<tr>\n<td>The translation was truncated<\/td>\n<td>Raise <strong>Max output tokens<\/strong>.<\/td>\n<\/tr>\n<tr>\n<td>&#8220;Failed to parse the model response as JSON&#8221;<\/td>\n<td>The model did not answer in JSON. Turn <strong>JSON mode<\/strong> on if it is off; if it is already on, the model is too weak for the job &#8211; try a larger one.<\/td>\n<\/tr>\n<tr>\n<td>An SSL or certificate error<\/td>\n<td>The endpoint needs a valid certificate. For a server on your own machine, use <code>http:\/\/<\/code> with a local address instead.<\/td>\n<\/tr>\n<tr>\n<td>Model does not exist<\/td>\n<td>The model name does not match what the endpoint serves.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Limits worth knowing<\/h2>\n<ul>\n<li><strong>Synchronous only.<\/strong> No Batch API, no Background mode, and therefore none of the 50% batch<br \/>\ndiscount.<\/li>\n<li><strong>Costs are what you type in.<\/strong> The plugin cannot read your provider&#8217;s prices, so the figures in<br \/>\nHistory are only as accurate as the two price fields.<\/li>\n<li><strong>One custom endpoint at a time<\/strong>, like every other engine.<\/li>\n<li><strong>No connection test.<\/strong> The first real translation is the test; run it on a short post.<\/li>\n<li><strong>Long model names are shortened in History.<\/strong> The full name is always sent to your endpoint, but<br \/>\nthe History page shows at most the first 60 characters of it.<\/li>\n<\/ul>\n<h2>Developer reference<\/h2>\n<p>The request body can be changed before it is sent &#8211; to add <code>temperature<\/code>, <code>top_p<\/code> or a vendor<br \/>\nextension the settings screen does not offer:<\/p>\n<pre><code class=\"language-php\">add_filter( 'latwmp_custom_openai_request_body', function ( $body, $model, $prompt ) {\n    $body['temperature'] = 0.2;\n    return $body;\n}, 10, 3 );<\/code><\/pre>\n<p>Whatever the filter returns is sent as-is. The engine&#8217;s provider id is <code>custom_openai<\/code>; its settings<br \/>\nlive in the shared <code>latwm_ai_settings<\/code> option under the <code>custom_openai_*<\/code> keys, and are removed when<br \/>\nthe PRO add-on is uninstalled.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Available in PRO (version 2.4.0 and later). Besides the built-in engines, the plugin can translate through any API that speaks OpenAI&#8217;s chat format. You give it\u2026<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_latwm_translated_slug":"","footnotes":""},"categories":[4],"tags":[],"class_list":["post-110","post","type-post","status-publish","format-standard","hentry","category-latw-multilingual"],"_links":{"self":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/posts\/110","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/comments?post=110"}],"version-history":[{"count":1,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/posts\/110\/revisions"}],"predecessor-version":[{"id":115,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/posts\/110\/revisions\/115"}],"wp:attachment":[{"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/media?parent=110"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/categories?post=110"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/docs.latw.ai\/pl\/wp-json\/wp\/v2\/tags?post=110"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}