> ## Documentation Index
> Fetch the complete documentation index at: https://docs.xquik.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Haystack Twitter search API & RAG Python guide

> Build Haystack RAG pipelines with Twitter search and user timeline APIs, typed Documents, citations, pagination, async Python, and exact error handling.

<blockquote className="agent-llms-directive">
  For the complete documentation index, see <a href="/llms.txt">llms.txt</a>.
</blockquote>

Haystack is an open source framework for Python AI applications. Compare frameworks for building RAG and agent pipelines. Use [`xquik-haystack`](https://github.com/Xquik-dev/xquik-haystack) for current tweets. This Haystack AI framework gives you typed tweet `Document` objects. Each includes text, authors, metrics, and URLs.

This Haystack AI API integration supports RAG pipelines and agent workflows. You can follow this guide when building AI applications with Haystack and current tweets. The integration never grants write access.

The integration provides 2 read-only components:

<CardGroup cols={2}>
  <Card title="Search public tweets" icon="search">
    `XquikTweetSearch` searches keywords, hashtags, accounts, conversations, and
    exact phrases.
  </Card>

  <Card title="Fetch user timelines" icon="user">
    `XquikUserTweetsFetcher` retrieves one public account's tweets and optional
    replies.
  </Card>
</CardGroup>

Use these components for search, timelines, monitoring, and retrieval-augmented generation. They can retrieve relevant tweets for Haystack AI agents and Haystack AI RAG pipelines.

Pipelines and agents can both use these components. Wrap `XquikTweetSearch` with `ComponentTool` when an agent needs a search tool called `search_current_tweets`.

Use the [followers API](/api-reference/x/followers) for follower exports. Use the [write API](/api-reference/x-write/create-tweet) for approved publishing. Those actions stay outside this integration.

## Install Haystack and Xquik

Use Python 3.10 or newer. Pin both packages for repeatable pipeline builds.

```bash theme={null}
python -m pip install "xquik-haystack==0.1.3" "haystack-ai==3.0.0"
```

Install inside a virtual environment. PyPI hosts release `0.1.3`. Haystack `3.0.0` supports sync and async runs.

The `pip install` command pins both packages for repeatable builds.

Create an [Xquik API key](/x-api-quickstart), then export it locally.

```bash theme={null}
export XQUIK_API_KEY="xq_..."
```

Never embed production keys in pipeline YAML, notebooks, or source control. Load the environment variable through a Haystack `Secret` object.

## Search Twitter tweets in Python

`XquikTweetSearch` calls the `GET /x/tweets/search` search endpoint. It accepts standard X search syntax and structured filters. Use it for each Twitter API keyword search or exact phrase query.

```python theme={null}
from haystack import Pipeline
from haystack.utils import Secret
from haystack_integrations.components.websearch.xquik import XquikTweetSearch

search = XquikTweetSearch(
    api_key=Secret.from_env_var("XQUIK_API_KEY"),
    top_k=20,
    query_type="Latest",
)

pipeline = Pipeline()
pipeline.add_component("twitter_search", search)

result = pipeline.run(
    {
        "twitter_search": {
            "query": '"retrieval augmented generation" lang:en -filter:retweets'
        }
    }
)

documents = result["twitter_search"]["documents"]
links = result["twitter_search"]["links"]
```

Use `Latest` for recent monitoring. Use `Top` for engagement-ranked discovery. Store tweet IDs because rankings can change.

Save every search query beside its tweet IDs. Set `top_k` to cap the number of tweets in each run.

### Build focused tweet searches

| Search intent | Query example |
| - | - |
| Exact phrase | `"retrieval augmented generation"` |
| Account posts | `from:deepset_ai haystack` |
| Hashtag search | `#haystack #rag` |
| Date window | `haystack since:2026-07-01 until:2026-08-01` |
| Exclude reposts | `haystack -filter:retweets` |

The API also supports structured search filters. Pass them through `extra_params` during initialization.

```python theme={null}
search = XquikTweetSearch(
    top_k=50,
    query_type="Latest",
    extra_params={
        "language": "en",
        "mediaType": "links",
        "minFaves": 10,
        "replies": "exclude",
    },
)
```

See the [tweet search API](/api-reference/x/search-tweets) for author, reply, quote, URL, and conversation filters. Keep timestamps and cursors outside `extra_params`.

## Choose a search window

Bound every retrieval with timestamps.

```python theme={null}
result = search.run(
    query="haystack agents",
    since_time="2026-07-31T00:00:00Z",
    until_time="2026-08-01T00:00:00Z",
)
```

Save each window. Deduplicate overlaps by `Document.meta["id"]`.

## Fetch a Twitter user timeline

Use `XquikUserTweetsFetcher` for `GET /x/users/{id}/tweets`. Pass a username or numeric X user ID.

```python theme={null}
from haystack.utils import Secret
from haystack_integrations.components.websearch.xquik import (
    XquikUserTweetsFetcher,
)

timeline = XquikUserTweetsFetcher(
    api_key=Secret.from_env_var("XQUIK_API_KEY"),
    include_replies=False,
    include_parent_tweet=False,
)

result = timeline.run(user_id="deepset_ai")
documents = result["documents"]
```

Set `include_replies=True` for replies. Enable parent tweets only when reply context matters. Choose search for many accounts and timelines for one account.

## Understand Haystack document fields

Each tweet becomes one Haystack `Document`. Tweet text becomes `Document.content`. Stable fields become metadata.

| Document field | Stored tweet value |
| - | - |
| `content` | Tweet text |
| `meta.endpoint` | Search or timeline route |
| `meta.id`, `meta.url` | Tweet ID and canonical URL |
| `meta.created_at`, `meta.lang` | Timestamp and language |
| `meta.conversation_id`, `meta.in_reply_to_*` | Conversation and parent fields |
| `meta.is_reply`, `meta.is_quote_status` | Reply and quote flags |
| `meta.like_count`, `meta.retweet_count`, `meta.reply_count` | Likes, reposts, and replies |
| `meta.quote_count`, `meta.view_count`, `meta.bookmark_count` | Quotes, views, and bookmarks |
| `meta.author` | Author ID, username, name, and verified flag |

Missing fields stay absent. Never treat missing metrics as 0. The `links` output contains each available `meta.url`.

## Keep evidence in RAG citations

Store tweet IDs and canonical URLs before embedding tweet text. This keeps evidence after ranking or joining.

```python theme={null}
citation_rows = []

for document in documents:
    tweet_id = document.meta.get("id")
    tweet_url = document.meta.get("url")
    if tweet_id and tweet_url:
        citation_rows.append(
            {
                "tweet_id": tweet_id,
                "url": tweet_url,
                "created_at": document.meta.get("created_at"),
                "author": document.meta.get("author"),
            }
        )
```

Require supplied URLs for citations. Reject URLs absent from retrieved documents. Keep `conversation_id` for reply threads.

### Treat tweet text as untrusted context

Tweets can contain prompt injection and unsafe URLs. Never treat tweet text as a system instruction. Keep tweets separate from instructions and tool permissions.

Limit retrieval by topic and time. Keep tweet IDs, authors, timestamps, and URLs. Require citations. Review sensitive conclusions.

Likes and reposts rank results. They do not prove accuracy.

Treat outputs from large language models (LLMs) as proposals, not evidence.

## Index tweets or retrieve them live

Choose the pipeline pattern that matches freshness requirements.

<CardGroup cols={3}>
  <Card title="Live retrieval" icon="radio">
    Search during each question for recent tweets and active events.
  </Card>

  <Card title="Indexed corpus" icon="database">
    Store embeddings for repeated research across stable windows.
  </Card>
</CardGroup>

Expect current results, not guaranteed real time delivery. Keep tweet records separate from embeddings. Use `meta.id` for deduplication.

## Paginate without duplicate tweets

Both components return `has_more` and `next_cursor`. Keep the request unchanged. Treat cursors as opaque strings.

```python theme={null}
search = XquikTweetSearch(top_k=100, query_type="Latest")
page = search.run(query="haystack ai")

documents_by_tweet_id = {}

while True:
    for document in page["documents"]:
        tweet_id = document.meta.get("id")
        if tweet_id:
            documents_by_tweet_id[tweet_id] = document

    if not page["has_more"] or not page["next_cursor"]:
        break

    page = search.run(
        query="haystack ai",
        cursor=page["next_cursor"],
    )

documents = list(documents_by_tweet_id.values())
```

Save each cursor after its documents. Never edit a cursor.

Deduplicate on `Document.meta["id"]`.

## Run Haystack pipelines asynchronously

Haystack 3 uses one `Pipeline` class. Both Xquik components expose `run_async()`.

```python theme={null}
import asyncio

from haystack import Pipeline
from haystack_integrations.components.websearch.xquik import XquikTweetSearch


async def search_tweets():
    pipeline = Pipeline()
    pipeline.add_component(
        "twitter_search",
        XquikTweetSearch(top_k=25, query_type="Latest"),
    )
    return await pipeline.run_async(
        {"twitter_search": {"query": "haystack agents"}}
    )


result = asyncio.run(search_tweets())
```

Use async runs in web servers. Use sync runs for scripts and scheduled jobs.

## Handle every documented error

The integration raises `httpx.HTTPStatusError`. Branch on the canonical status before retrying.

| Status | Meaning | Action |
| - | - | - |
| `400` | Invalid request | Fix it. Do not retry unchanged. |
| `401` | Invalid API key | Add a valid API key. |
| `402` | Account action needed | Check the account balance. |
| `404` | Timeline user not found | Check the username or user ID. |
| `424` | X retrieval dependency failed | Retry with bounded backoff. |
| `429` | Rate limit exceeded | Back off before retrying. |
| `502` | X retrieval error | Retry up to the configured limit. |

```python theme={null}
import httpx

from haystack_integrations.components.websearch.xquik import XquikTweetSearch

search = XquikTweetSearch(top_k=20, max_retries=3)

try:
    result = search.run(query="haystack ai")
except httpx.HTTPStatusError as error:
    status = error.response.status_code
    if status in {400, 401, 402, 404}:
        raise RuntimeError(f"Fix Xquik request before retrying: HTTP {status}") from error
    if status in {424, 429, 502}:
        raise RuntimeError(f"Retry Xquik request with backoff: HTTP {status}") from error
    raise
```

Record status codes, never credentials. Cap retries to prevent unbounded agent loops.

## Pipeline handoff

Use this shape when Haystack hands results to a vector store, evaluation job, queue, CSV export, or dashboard.

<CardGroup cols={2}>
  <Card title="Document rows" icon="file-text">
    Store content, tweet IDs, URLs, timestamps, authors, and public metrics.
  </Card>

  <Card title="Citation links" icon="link">
    Join each canonical URL to `meta.id`.
  </Card>

  <Card title="Pagination checkpoint" icon="list-tree">
    Store the request, options, `has_more`, and `next_cursor`.
  </Card>

  <Card title="Failure branch" icon="route">
    Store each HTTP status with its pipeline run ID.
  </Card>
</CardGroup>

Keep handoff records separate from embeddings. Later runs can refresh tweets without rebuilding the pipeline.

## Haystack component or direct REST API

Does your pipeline already return `Document` objects? Add these components directly. They normalize tweet text, metadata, URLs, and pagination.

Use direct Xquik REST routes for followers, following, replies, quotes, reposts, lists, communities, media, trends, monitors, approved writes, and extra query options.

Both approaches use the same Xquik contracts. Components only support tweet search and user timelines.

## Migrate to Haystack 3

Haystack 3 replaces `AsyncPipeline` with `Pipeline`. Call `await pipeline.run_async(...)` for concurrent execution. Keep `meta.id` as the tweet identity.

Test with `haystack-ai==3.0.0` before upgrading. Check the [official migration guide](https://docs.haystack.deepset.ai/docs/migration) for other changes.

## Common Haystack Twitter API questions

### What is Haystack AI?

Haystack links search, RAG, and AI agents inside Python pipelines.

### How does the Twitter API search tweets?

It accepts queries, filters, ordering, and cursors. Xquik returns tweets and citation URLs.

### What is a Twitter search API Python workflow?

Install both packages. Run `XquikTweetSearch`, then process documents, links, and cursors.

### How does an API Twitter search workflow keep cursors?

Save `next_cursor` after each page. Reuse every filter.

### How does the Twitter API search tweets by keyword?

Pass keywords, phrases, hashtags, accounts, or filters. Choose `Latest` or `Top`.

### Haystack AI vs LangChain: which fits Twitter RAG?

Choose Haystack for `Document` pipelines. Choose LangChain for its retrievers and tools.

### Where is the Haystack AI GitHub integration?

The [Xquik Haystack repository](https://github.com/Xquik-dev/xquik-haystack) contains the package and offline tests.

### Can a Haystack AI agent publish tweets?

No. This integration reads searches and timelines. Use the write API for approved tweets.

### Does this require a Haystack Enterprise Platform?

No. Run the open-source components in your Haystack setup.

## Source and contracts

* [Xquik Haystack repository](https://github.com/Xquik-dev/xquik-haystack)
* [Xquik Haystack PyPI package](https://pypi.org/project/xquik-haystack/)
* [Haystack 3.0 release](https://github.com/deepset-ai/haystack/releases/tag/v3.0.0)
* [Search Tweets API](/api-reference/x/search-tweets)
* [Get User Timeline API](/api-reference/x/user-tweets)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.