Skip to main content

TL;DR

  • Learn how to use Orq AI Gateway
  • Connect primary and fallback AI providers to avoid vendor lock-in
  • Enable streaming for real-time responses and better UX
  • Add a knowledge base with custom docs for contextual answers
  • Set up caching for recurring requests
  • Build a production-ready customer support agent in minutes

Overview

This tutorial builds a customer support application in Node.js using AI Gateway, where support queries have access to relevant business context from a Knowledge Base. The system includes a primary model (GPT-4o) and a fallback model (Claude Sonnet) that automatically activates during rate limits or outages. The tutorial also covers caching for user queries, identity tracing to monitor per-user LLM request volumes, and Thread tracking to visualize complete conversation flows between users and the assistant.

What is AI gateway?

AI Gateway is a single unified API endpoint that lets you seamlessly route and manage requests across multiple AI model providers (e.g., OpenAI, Anthropic, Google, AWS). This functionality comes in handy, when you want to:
  • Avoid dependency on a single provider (vendor lock-in)
  • Automatically switch between providers in case of an outage
  • Scale reliably when the usage surges

Build the customer support chat

1

Set up the Node.js project

Inside the IDE of choice, set up the Node.js project. This tutorial uses npm; alternatives such as pnpm are also supported.
First, inside Orq dashboard create a project that we can assign API keys to by clicking the + button next to Project menu:Add projectCreate a new project named CustomerSupportNew project named CustomerSupport created in the Orq.ai dashboardTo find the API key, navigate to Settings > Organization > API Keys and copy the key.API Keys management table listing keys with columns for name, type, status, permissions, and created by.From the dropdown, select the CustomerSupport project to assign the API key to:Dropdown showing CustomerSupport project selected for API key assignmentCreate a .env file with the following content, replacing the placeholder with the actual API key:
Add .env to your .gitignore
Create the customer-support.ts file with a Hello World example:
customer-support.ts
To execute the file from the terminal run:
Hello world
2

Streaming data in real time

This step uses the OpenAI gpt-4o model to generate responses. To connect any other model such as claude-sonnet-4-6, follow the same steps. To enable models in AI Gateway:
  1. Navigate to Integrations
  2. Select OpenAI
  3. Click on View integration
Streaming dataClick on Setup your own API keySet up APILog in to OpenAI’s API platform and copy your secret key:OpenAINavigate back to the Orq.ai dashboard and paste the API keys inside the pop-up window that appears after clicking the Setup your own API key buttonOpenAI setupBy default, when you make a POST request, the connection remains open until the entire response is ready, and then it closes.However, when you use streaming, the API switches to a Server-Sent Events (SSE) connection. This keeps the HTTP connection open and sends the response in small, real-time chunks as the data becomes available and is essential for real-time customer chat interactions.
customer-support.ts
Streaming is ideal for applications that display text as it is generated, such as chat interfaces or live assistants, improving perceived responsiveness:
3

Retries & fallbacks

Orq.ai allows automatic fallback to alternative models if the primary fails. If gpt-4o hits a rate limit or downtime, the request automatically retries and may fall back to Anthropic claude-sonnet-4-6 or gpt-5-mini. Make sure the models are enabled in Orq.ai.
4

Caching

Orq.ai supports response caching to reduce latency and API usage for repeated requests. It uses exact_match caching, where the cache key is generated from the exact model, messages, and all parameters, ensuring identical requests hit the cache. The TTL (time-to-live) specifies how long the response is cached (e.g., 3600 seconds for 1 hour, max 86400 seconds). Below is a TypeScript implementation with caching, retries, and fallbacks:
On the first run, the request shows cache-miss inside Traces.Cache missThe cache is stored after the command runs for the first time. The reason for cache-miss on the first run is that Orq.ai has no prior response stored for that exact cache key. Read more about cache here.Running the same request a second time within the TTL shows cache-hit inside Traces, meaning Orq.ai retrieved the cached response.Cache hit
5

Knowledge Base

When to use:
  • When you want to enhance a foundational model’s responses with custom, domain-specific knowledge using Retrieval-Augmented Generation (RAG).
  • Orq.ai’s built-in RAG feature enables creation of a Knowledge Base from documents (e.g., FAQs, manuals, or PDFs)
  • When you want to add a Vector Database (e.g., Pinecone, Qdrant) for control over embeddings and retrieval. For more see Using Vector databases with Orq
Orq.ai Knowledge Bases support the following file types: pdf, txt, docx, csv, xls (10 MB max). Encrypted files are not supported.The following parameters control Knowledge Base creation:Run the code to create a Knowledge Base:
customer-support.ts
This is how a successful response should look like:
Save the Knowledge Base ID _id as YOUR_KNOWLEDGE_ID in the .env file, replacing the placeholder with the actual value from the response above:
To complete this step with the GUI, see Create a Knowledge Base.
6

Add files to the Knowledge Base

Inside the main repository create a documents directory and place the documents to upload there. Orq.ai supports document types: pdf, txt, docx, csv, xls (10 MB max).Run the following code to upload the documents:
customer-support.ts
This is how a successful response should look like:
Add the file ID _id to the .env file, replacing the placeholder with the actual value from the response above:
To complete this step with the GUI, see Upload a file.
7

Connect the files with the Knowledge Base as datasource

This is how a successful response looks like:
Confirm YOUR_KNOWLEDGE_ID is present in .env from the previous step.The uploaded file is now visible under the Knowledge Base:Uploaded file visible under the Knowledge Base in the Orq.ai dashboardTo complete this step with the GUI, see Creating a new Datasource.When documents are uploaded to a Knowledge Base, Orq.ai breaks them into smaller pieces of text called chunks. Think of it like dividing a book into manageable paragraphs or sections rather than trying to process the entire book at once.ChunksThis is the customer support chat with connected Knowledge Base:
After running the code, the Knowledge Base retrieval is visible on the Orq.ai dashboard.Traces
8

Identity Tracking

When to use:
  • You want to identify and remember the user between chats or sessions.
  • You need to audit who asked what (e.g., Alice Smith asked about “refunds”).
  • You’re building user profiles, dashboards, or integrating with a CRM (e.g., Salesforce, HubSpot).
  • When the application involves external B2B clients and monitoring call volume and cost per client is required
For more details see Identity TrackingWhen prototyping with cURL, paste the code snippet with YOUR_API_KEY, YOUR_IDENTITY_ID and YOUR_DEPLOYMENT_KEY variables:
After the code snippet runs successfully, the number of requests sent by the selected Identity is visible under Identity Analytics. See also budget control.Control the budget
9

Thread tracking

When to use:
  • Understand the back-and-forth between the user and the assistant
  • Track context drift in long conversations
  • Make sense of multi-step conversations at a glance
To enable Thread tracking, use this version of the customer support app. To learn more, see Threads.
After the code snippet runs successfully, a detailed breakdown of the API call is visible under Traces > Threads.Thread breakdown visible under Traces in the Orq.ai dashboardSending a request again with the same thread.id (support-TICKET-789-<timestamp>) for both initial and follow-up requests groups them in the same Thread:Two requests grouped under the same Thread ID in the Orq.ai Traces view
10

Dynamic Inputs

When to use:

Advanced framework integrations

Orq.ai’s AI Gateway integrates with popular AI development frameworks, allowing existing tools and workflows to benefit from gateway features like fallbacks, caching, and observability.

LangChain Integration

Orq.ai works natively with LangChain by simply pointing to the AI Gateway endpoint. This gives access to fallback models, caching, and Knowledge Base retrieval while using LangChain’s abstractions. For a more detailed guide, see LangChain integration.

DSPy

DSPy programs can route through Orq.ai to gain automatic prompt optimization alongside gateway reliability features. For a more detailed guide, see DSPy Integration.

Base URL configuration

Conclusion

Orq.ai’s AI Gateway provides a unified, scalable, and production-ready solution for building reliable AI applications. By routing through a single API endpoint, the application gains:
  1. Unified access: Connect to multiple AI providers (OpenAI, Anthropic, AWS) through one API
  2. High availability: Automatic fallbacks and retries ensure the application stays online
  3. Cost efficiency: Response caching reduces API costs and latency
  4. Smart context: Built-in Knowledge Base integration for domain-specific answers
  5. Production observability: Comprehensive Traces and OTEL compatibility for monitoring
  6. Flexible deployment: Cloud, on-premises, or edge options to meet deployment needs