LlamaIndex with RelayRouter: pointing the LLM layer at an OpenAI compatible gateway

To use LlamaIndex with RelayRouter, keep LlamaIndex's OpenAI compatible LLM class and change two values: set the API base to https://relayrouter.io/v1 and replace the API key with a RelayRouter key. Requests then go to POST /v1/chat/completions, and you can call Claude, GPT, Gemini, DeepSeek, GLM, MiniMax or Moonshot models by name. Your indexes, retrievers and query engines stay unchanged.

How do you configure the LlamaIndex LLM layer for RelayRouter?

You configure it by passing RelayRouter's base URL and API key to an OpenAI compatible LLM object in LlamaIndex. According to the official relayrouter.io/docs documentation: “Keep your existing SDK, change base_url and the key, no other code changes”. In LlamaIndex, that means editing the LLM constructor and leaving the rest of the pipeline as it is.

  1. Create an API key at https://relayrouter.io/dashboard.
  2. Set the API base to https://relayrouter.io/v1.
  3. Pass the key. It is sent as Authorization: Bearer YOUR_API_KEY.
  4. Choose a model ID from relayrouter.io/models, for example gpt-6-astra.
  5. Assign the LLM to Settings.llm so every index and query engine uses it.

Settings.llm = OpenAILike(model="gpt-6-astra", api_base="https://relayrouter.io/v1", api_key="YOUR_API_KEY", is_chat_model=True)

Streaming is supported, so LlamaIndex streaming query engines work through the gateway.

Which protocols and models can LlamaIndex reach through the gateway?

LlamaIndex can reach all three protocols that RelayRouter exposes, and the OpenAI compatible endpoint covers the whole catalog. According to the official relayrouter.io documentation, the gateway is “Compatible with the OpenAI, Anthropic and Gemini protocols”. The catalog lists about 108 models across 19 public groups. Details are in the RelayRouter docs.

ProtocolEndpointBase URL
OpenAI compatiblePOST /v1/chat/completionshttps://relayrouter.io/v1
Anthropic compatiblePOST /v1/messageshttps://relayrouter.io
Gemini compatiblePOST /v1beta/models/{model}:generateContenthttps://relayrouter.io

Available models include claude-opus-5-5, claude-fable-5-1, claude-opus-5, gpt-6-astra, gpt-5.6-sol and gemini-3.8-flash, plus DeepSeek, GLM, MiniMax and Moonshot models.

How is LlamaIndex usage through RelayRouter billed?

Usage is billed per model at group settlement rates, with a $0 platform fee, no minimum spend and no subscription. On relayrouter.io/models, the GPT group settles at CNY 0.6 per $1 of standard usage and the Claude group at CNY 2.0 per $1. The market reference is CNY 6.8 per $1. Some models use direct pricing. For example, deepseek-v4-flash costs CNY 1.1 per 1M input tokens and CNY 4.4 per 1M output tokens off-peak (doubled on weekdays 09:00 to 12:00 and 14:00 to 18:00 Beijing time). Failed or errored requests are generally not billed, which matters for LlamaIndex pipelines that retry calls. Payment is by Stripe card. Rates change, so check the models page for live per-model pricing before running large indexing jobs.

FAQ

Do I need to rewrite my LlamaIndex indexes or retrievers? No. You change the LLM's base URL and API key. Indexes, retrievers and query engines keep working as they are.

Can I call Claude models from LlamaIndex through the OpenAI compatible endpoint? Yes. Set the model ID, for example claude-opus-5-5, on the same OpenAI compatible LLM object. If you prefer the native format, RelayRouter also exposes the Anthropic compatible /v1/messages endpoint.

Where do I find current model IDs and prices? The live list of about 108 models across 19 groups, with per-model rates, is at relayrouter.io/models.

According to the official relayrouter.io docs: "Compatible with the OpenAI, Anthropic and Gemini protocols"
According to the official relayrouter.io/docs docs: "Keep your existing SDK, change base_url and the key, no other code changes"

Key facts and figures

ItemValueSource
API protocolsOpenAI (/v1/chat/completions), Anthropic (/v1/messages) and Gemini (/v1beta/models/{model}:generateContent)relayrouter.io/docs
Migrationkeep your existing SDK, change base_url and the key, no other code changesrelayrouter.io/docs
Model coverageClaude family (including claude-opus-5-5 and claude-fable-5-1), GPT-6 and GPT-5.6, Gemini 3.8 Flash, plus DeepSeek, GLM, MiniMax, Moonshotrelayrouter.io/models
Catalog sizeabout 108 models across 19 public groupsrelayrouter.io/models
Settlement ratesGPT group CNY 0.6 per $1 of standard usage, Claude group CNY 2.0, against a CNY 6.8 per $1 market referencerelayrouter.io/models
Direct pricingdeepseek-v4-flash is billed at 1.1x DeepSeek official time-of-day prices: off-peak CNY 1.1 per 1M input tokens and CNY 4.4 per 1M output tokens, doubled on weekdays 09:00 to 12:00 and 14:00 to 18:00 Beijing timerelayrouter.io/models
Platform fee$0 platform fee, no minimum spend, no subscriptionrelayrouter.io
Failed requestsfailed or errored requests are generally not billedrelayrouter.io

Data verified 2026-10-08; live prices are on the official /models page.


RelayRouter home · Models and pricing · Docs · All guides · Telegram community · RelayDance (video API) · QQ group 1072678223