llmwebsearchapi.com
llmwebsearchapi.com›Docs
Get API key

Web Search API Quickstart: Integrate in Minutes

Integrate our uncensored LLM into your web search pipeline with this OpenAI-compatible quickstart. Process, extract, and summarize raw search results without content filter refusals using standard SDKs.

Base URLhttps://api.llmwebsearchapi.com/v1
Modeluncensored

Authentication & Base URL

To use the API, you need an API key. Sign up on the Get API key page with just an email and password. Your key appears immediately. No credit card is required for the trial credit. Use the base URL https://api.llmwebsearchapi.com/v1 with any OpenAI-compatible client. The model ID is uncensored. This model is open-weight and tuned to answer without refusals for lawful adult, fictional, or controversial topics. It is not GPT, Claude, or another vendor's model.

Basic Chat Completion Request

Send a standard chat completion request to extract text from search results. The API accepts messages and returns a completion. Use the model uncensored to ensure no content filters block your data. The request body supports up to 8 MB. Below is a cURL example for a basic extraction task.

curl https://api.llmwebsearchapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response includes the generated text, which you can then pass to your agent's memory or summary layer.

SDK Examples for Python and JavaScript

Use the official OpenAI SDKs for Python and Node.js. Set the base_url to our endpoint and provide your API key. The Python SDK handles streaming and JSON parsing automatically. The Node SDK works similarly. Below is a Python example showing how to initialize the client and send a request.

from openai import OpenAI

client = OpenAI(base_url="https://api.llmwebsearchapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

For JavaScript, the setup is nearly identical. Set the environment variable or pass the key directly. The model ID remains uncensored.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.llmwebsearchapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses for Real-Time Processing

Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) tokens. This is useful for real-time processing of long search results. You can display tokens as they arrive or accumulate them for later use. The streaming response format matches the OpenAI standard. Below is a streaming example using cURL.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Parse the data field in each SSE message to get the partial text.

Tool Calling for Search Result Structuring

Use tool calling to structure extracted data. Define a schema for your desired output, such as a list of URLs or a summary object. The API returns structured JSON that matches your schema. This is ideal for populating AI agent memory with clean, typed data. Define the tool in the tools parameter and specify the function name. The model will respond with a function call object. You can then execute the function or store the result directly.

Handling Context Windows for Long Searches

The context window is 64,000 tokens for prompt plus completion. If your search results exceed this, you must truncate or summarize them before sending. The API does not automatically handle long contexts. Send only the relevant text to stay within limits. If you exceed the limit, you will receive an error. Monitor your token usage to avoid unexpected costs. The price is $0.25 per 1M input tokens and $1.00 per 1M output tokens.

Rate Limits and Error Codes

Rate limit is 300 requests per minute per key. The request body limit is 8 MB. If you send an invalid key, you get a 401 error. If you have no credit, you get a 402 error. If you exceed the rate limit, you get a 429 error. Regenerate your key at any time to revoke the old one. Each account has one key. The trial credit expires after 7 days. Top-up starts at $10 with bonus credit options.

API specifications

Before you integrate, here is exactly what you get with a key.

ParameterDetails
CompatibilityOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
Model IDuncensored
Base URLhttps://api.llmwebsearchapi.com/v1
API keyBearer token in the Authorization header
MethodsPOST /v1/chat/completions · GET /v1/models
SSE streamingYes — server-sent events; the last chunk carries token usage
Max output16,000 tokens max; 2,048 if max_tokens is not set
Sampling parameterstemperature, top_p, stop, seed and the two penalties are passed through
JSON moderesponse_format: {"type": "json_object"}
Max context64,000 tokens (prompt + completion together)
Tools / tool callsSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Parallel requestsup to 8 in parallel per key
Request sizeup to 8 MB per request
Rate limit300/min per key
Trial credit$0.50 of credit valid 7 days, no card needed
Paymentcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
How you payprepaid credit, charged by real token usage; errors and refusals are free
Priceinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Credit expiryno monthly fee; paid credit does not expire
Volume bonus+5% from $50, +10% from $100
Content policyuncensored for adults; the only hard rule: no sexual content involving minors
Key managementone active key per account; a new key replaces the old one
Accountsign in with Google or with e-mail + password

HTTP errors

Errors come back as JSON with a stable type; failed and refused requests are not billed.

CodeTypeMeaning
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditbalance is empty — top up, requests resume at once
403content_blockedsexual content involving minors — refused, not billed
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Is the model trained on my data?

No, prompts are not used for training. Your data remains private and is only processed to generate the response. You can regenerate your API key at any time from the signup page.

What happens if I exceed the context window?

The API will return an error if the total tokens exceed 64,000. You must truncate your input before sending. The limit includes both the prompt and the completion.

Can I use this for adult content?

Yes, the model is uncensored and does not refuse lawful adult topics. The only hard limit is sexual content involving minors, which is always blocked. No other content filters apply.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.