Web Search API Quickstart: Integrate in Minutes
Integrate our uncensored LLM into your web search pipeline with this OpenAI-compatible quickstart. Process, extract, and summarize raw search results without content filter refusals using standard SDKs.
https://api.llmwebsearchapi.com/v1uncensoredAuthentication & Base URL
To use the API, you need an API key. Sign up on the Get API key page with just an email and password. Your key appears immediately. No credit card is required for the trial credit. Use the base URL https://api.llmwebsearchapi.com/v1 with any OpenAI-compatible client. The model ID is uncensored. This model is open-weight and tuned to answer without refusals for lawful adult, fictional, or controversial topics. It is not GPT, Claude, or another vendor's model.
Basic Chat Completion Request
Send a standard chat completion request to extract text from search results. The API accepts messages and returns a completion. Use the model uncensored to ensure no content filters block your data. The request body supports up to 8 MB. Below is a cURL example for a basic extraction task.
curl https://api.llmwebsearchapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'The response includes the generated text, which you can then pass to your agent's memory or summary layer.
SDK Examples for Python and JavaScript
Use the official OpenAI SDKs for Python and Node.js. Set the base_url to our endpoint and provide your API key. The Python SDK handles streaming and JSON parsing automatically. The Node SDK works similarly. Below is a Python example showing how to initialize the client and send a request.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmwebsearchapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)For JavaScript, the setup is nearly identical. Set the environment variable or pass the key directly. The model ID remains uncensored.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmwebsearchapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses for Real-Time Processing
Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) tokens. This is useful for real-time processing of long search results. You can display tokens as they arrive or accumulate them for later use. The streaming response format matches the OpenAI standard. Below is a streaming example using cURL.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Parse the data field in each SSE message to get the partial text.
Tool Calling for Search Result Structuring
Use tool calling to structure extracted data. Define a schema for your desired output, such as a list of URLs or a summary object. The API returns structured JSON that matches your schema. This is ideal for populating AI agent memory with clean, typed data. Define the tool in the tools parameter and specify the function name. The model will respond with a function call object. You can then execute the function or store the result directly.
Handling Context Windows for Long Searches
The context window is 64,000 tokens for prompt plus completion. If your search results exceed this, you must truncate or summarize them before sending. The API does not automatically handle long contexts. Send only the relevant text to stay within limits. If you exceed the limit, you will receive an error. Monitor your token usage to avoid unexpected costs. The price is $0.25 per 1M input tokens and $1.00 per 1M output tokens.
Rate Limits and Error Codes
Rate limit is 300 requests per minute per key. The request body limit is 8 MB. If you send an invalid key, you get a 401 error. If you have no credit, you get a 402 error. If you exceed the rate limit, you get a 429 error. Regenerate your key at any time to revoke the old one. Each account has one key. The trial credit expires after 7 days. Top-up starts at $10 with bonus credit options.
API specifications
Before you integrate, here is exactly what you get with a key.
| Parameter | Details |
|---|---|
| Compatibility | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Model ID | uncensored |
| Base URL | https://api.llmwebsearchapi.com/v1 |
| API key | Bearer token in the Authorization header |
| Methods | POST /v1/chat/completions · GET /v1/models |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| JSON mode | response_format: {"type": "json_object"} |
| Max context | 64,000 tokens (prompt + completion together) |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Parallel requests | up to 8 in parallel per key |
| Request size | up to 8 MB per request |
| Rate limit | 300/min per key |
| Trial credit | $0.50 of credit valid 7 days, no card needed |
| Payment | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Price | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Credit expiry | no monthly fee; paid credit does not expire |
| Volume bonus | +5% from $50, +10% from $100 |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
| Key management | one active key per account; a new key replaces the old one |
| Account | sign in with Google or with e-mail + password |
HTTP errors
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Is the model trained on my data?
No, prompts are not used for training. Your data remains private and is only processed to generate the response. You can regenerate your API key at any time from the signup page.
What happens if I exceed the context window?
The API will return an error if the total tokens exceed 64,000. You must truncate your input before sending. The limit includes both the prompt and the completion.
Can I use this for adult content?
Yes, the model is uncensored and does not refuse lawful adult topics. The only hard limit is sexual content involving minors, which is always blocked. No other content filters apply.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.