{
  "run_id": "2026-09-04-r1",
  "tool": "firecrawl",
  "mode": "search",
  "query_id": "Q51",
  "query_text": "Olostep total index size in pages",
  "input_file": 3,
  "timestamp_utc": "2026-09-04T09:27:36Z",
  "region": "ap-south-1",
  "latency_ms": 18617.4,
  "http_status": 200,
  "error": null,
  "credits_reported": 12,
  "tokens_reported": null,
  "results": [
    {
      "rank": 1,
      "url": "https://www.olostep.com/",
      "title": "Olostep: Web Data Infrastructure for AI Agents",
      "content": "Click to try\n\nWait...\n\n![](https://www.olostep.com/images/close-x-svgrepo-com-3.svg)\n\nOlostep Manifesto \\| [read more](https://www.olostep.com/blog/about-olostep) \u2192\n\n![](https://www.olostep.com/images/close-x-svgrepo-com-3.svg)\n\n[![](https://www.olostep.com/images/olostep-logo-cropped.svg)](https://www.olostep.com/)\n\n# Web Data Infrastructure for AI\n\nBuilt to power the Web's second user, Olostep is the best web agentic search, scraping and crawling API for AI\n\n[Start for free](https://www.olostep.com/auth) [Contact Sales](https://www.olostep.com/contact-sales)\n\nThank you! Your submission has been received!\n\nOops! Something went wrong while submitting the form.\n\n[Scrape](https://www.olostep.com/#w-tabs-0-data-w-pane-0) [Crawl](https://www.olostep.com/#w-tabs-0-data-w-pane-1) [Map](https://www.olostep.com/#w-tabs-0-data-w-pane-2) [Search](https://www.olostep.com/#w-tabs-0-data-w-pane-3) [Answer](https://www.olostep.com/#w-tabs-0-data-w-pane-4)\n\n![](https://www.olostep.com/images/svgexport-11-1.svg)\n\n## Trusted by the best startups **startups** in the world\n\n![](https://www.olostep.com/images/455e150089b14aedb083b23c8e8f157f__1_-removebg-preview.png)![](https://www.olostep.com/images/airops.png)![](https://www.olostep.com/images/podqi-logo.png)![](https://www.olostep.com/images/khoj_original-removebg-preview.png)![](https://www.olostep.com/images/svgexport-1-1.svg)![](https://www.olostep.com/images/finny_ai-removebg-preview.png)![](https://www.olostep.com/images/Logo_Contents_2025_Blue-scaled.png)![](https://www.olostep.com/images/athenahq-logo-black.png)![](https://www.olostep.com/images/CivilGrid_Logo-removebg-preview.png)![](https://www.olostep.com/images/logo.svg)![](https://www.olostep.com/images/plots_black.png)![](https://www.olostep.com/images/use-bear.avif)![](https://www.olostep.com/images/Searchable.svg)![](https://www.olostep.com/images/uman-logo.svg)![](https://www.olostep.com/images/verisave-logo.png)![](https://www.olostep.com/images/relay-app-image-removebg-preview.png)![](https://www.olostep.com/images/openmart_originak-removebg-preview.png)![](https://www.olostep.com/images/profound_logo-removebg-preview.png)![](https://www.olostep.com/images/centralize-logo.png)\n\n![](https://www.olostep.com/images/455e150089b14aedb083b23c8e8f157f__1_-removebg-preview.png)![](https://www.olostep.com/images/airops.png)![](https://www.olostep.com/images/podqi-logo.png)![](https://www.olostep.com/images/khoj_original-removebg-preview.png)![](https://www.olostep.com/images/svgexport-1-1.svg)![](https://www.olostep.com/images/finny_ai-removebg-preview.png)![](https://www.olostep.com/images/Logo_Contents_2025_Blue-scaled.png)![](https://www.olostep.com/images/athenahq-logo-black.png)![](https://www.olostep.com/images/CivilGrid_Logo-removebg-preview.png)![](https://www.olostep.com/images/logo.svg)![](https://www.olostep.com/images/plots_black.png)![](https://www.olostep.com/images/use-bear.avif)![](https://www.olostep.com/images/Searchable.svg)![](https://www.olostep.com/images/uman-logo.svg)![](https://www.olostep.com/images/verisave-logo.png)![](https://www.olostep.com/images/relay-app-image-removebg-preview.png)![](https://www.olostep.com/images/openmart_originak-removebg-preview.png)![](https://www.olostep.com/images/profound_logo-removebg-preview.png)![](https://www.olostep.com/images/centralize-logo.png)\n\n## One **API** to Automate Web Data\n\nSearch, scrape, structure and monitor the whole web with one\n\nAPI. Reliable, cost-effective, scalable. Handling Billions of requests\n\n[**Monitor** \\\\\n\\\\\nSet up monitors for events happening across the Web \\\\\n\\\\\nMonitor when AirOps publishes a new blog\\\\\n\\\\\nNew blog: The Next Chapter for AirOps...](https://www.olostep.com/dashboard/research-agents)\n\n[Latest updates SpaceX...\\\\\n\\\\\nUpdates - SpaceX...\\\\\n\\\\\n**Search** \\\\\n\\\\\nSearch the web with natural language and get ranked links](https://www.olostep.com/dashboard/playground) [Roosevelt's quote of critics\\\\\n\\\\\n\"It's not the critic who \\\\\n\\\\\ncounts; not the man\"\\\\\n\\\\\n**Answer** \\\\\n\\\\\nSearch the web and get AI-powered answers](https://www.olostep.com/dashboard/playground)\n\n[Olostep - The way we're gong to ratchet up our species\\\\\n\\\\\n**Scrape** \\\\\n\\\\\nGet real-time data from websites. Clean Markdown, HTML, Screenshots, JSON...](https://www.olostep.com/dashboard/playground) [Stripe Docs\\\\\n\\\\\n**Crawl** \\\\\n\\\\\nRetrieve all pages on a site and get their contents](https://www.olostep.com/dashboard/playground) [**Batch** \\\\\n\\\\\nProcess up to 100k URLs in 5-7 minutes](https://www.olostep.com/dashboard/playground)\n\n## Built for Developers\n\nObject-oriented API, native Python and NodeJS SDK clients,\n\nmetadata support, webhook events, easy to try and easy to scale\n\n[![](https://www.olostep.com/images/search-window-2.svg)\\\\\n\\\\\n/scrapes](https://www.olostep.com/#w-tabs-1-data-w-pane-0) [![](https://www.olostep.com/images/internet.svg)\\\\\n\\\\\n/crawls](https://www.olostep.com/#w-tabs-1-data-w-pane-1) [![](https://www.olostep.com/images/map.svg)\\\\\n\\\\\n/maps](https://www.olostep.com/#w-tabs-1-data-w-pane-2) [![](https://www.olostep.com/images/apple-shortcuts-1.svg)\\\\\n\\\\\n/batches](https://www.olostep.com/#w-tabs-1-data-w-pane-3) [![](https://www.olostep.com/images/search-1.svg)\\\\\n\\\\\n/searches](https://www.olostep.com/#w-tabs-1-data-w-pane-4) [![](https://www.olostep.com/images/bubble-search.svg)\\\\\n\\\\\n/answers](https://www.olostep.com/#w-tabs-1-data-w-pane-5) [![](https://www.olostep.com/images/bell-notification.svg)\\\\\n\\\\\n/monitors](https://www.olostep.com/#w-tabs-1-data-w-pane-6)\n\nGet clean data from any URL\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-2-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-2-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-2-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6result = client.scrapes.create(\n7    url_to_scrape=\"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n8    formats=[\"markdown\", \"html\"],\n9)\n10\n11print(result.markdown_content)\n12print(result.html_content)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const result = await client.scrapes.create({\n7  url: 'https://en.wikipedia.org/wiki/Alexander_the_Great',\n8  formats: ['markdown', 'html'],\n9})\n10\n11console.log(result.markdown_content)\n12console.log(result.html_content)\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/scrapes\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"url_to_scrape\": \"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n6    \"formats\": [\"markdown\", \"html\"]\n7  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/scrapes)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nCrawl all the subpages\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-3-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNodeJS](https://www.olostep.com/#w-tabs-3-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-3-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6crawl = client.crawls.create(\n7    start_url=\"https://olostep.com\",\n8    max_pages=100,\n9    include_urls=[\"/**\"],\n10    exclude_urls=[\"/collections/**\"],\n11    include_external=False,\n12)\n13\n14print(crawl.id, crawl.status)\n15\n16# Wait for completion and iterate pages\n17for page in crawl.pages():\n18    print(page.url)\n19    content = page.retrieve([\"markdown\"])\n20    print(content.markdown_content[:200])\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const crawl = await client.crawls.create({\n7  url: 'https://olostep.com',\n8  maxPages: 100,\n9  includeUrls: ['/**'],\n10  excludeUrls: ['/collections/**'],\n11  includeExternal: false,\n12})\n13\n14console.log(crawl.id, crawl.status)\n15\n16// Wait for completion and iterate pages\n17for await (const page of crawl.pages()) {\n18  console.log(page.url)\n19  const content = await client.retrieve({ retrieveId: page.retrieve_id, formats: ['markdown'] })\n20  console.log(content.markdown_content.slice(0, 200))\n21}\n```\n\n```bash\n1# Start crawl\n2curl -s -X POST \"https://api.olostep.com/v1/crawls\" \\\n3  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n4  -H \"Content-Type: application/json\" \\\n5  -d '{\n6    \"start_url\": \"https://olostep.com\",\n7    \"max_pages\": 100,\n8    \"include_urls\": [\"/**\"],\n9    \"exclude_urls\": [\"/collections/**\"],\n10    \"include_external\": false\n11  }'\n12\n13# Check status (replace <CRAWL_ID>)\n14curl -s \"https://api.olostep.com/v1/crawls/<CRAWL_ID>\" \\\n15  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n16\n17# Get pages (replace <CRAWL_ID>)\n18curl -s \"https://api.olostep.com/v1/crawls/<CRAWL_ID>/pages\" \\\n19  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n20\n21# Retrieve content (replace <RETRIEVE_ID>)\n22curl -s -G \"https://api.olostep.com/v1/retrieve\" \\\n23  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n24  --data-urlencode \"retrieve_id=<RETRIEVE_ID>\" \\\n25  --data-urlencode \"formats=markdown\"\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/crawls)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nGet all the URLs on a website\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-4-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-4-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-4-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6sitemap = client.maps.create(\n7    url=\"https://docs.olostep.com\",\n8    include_urls=[\"/features/**\"],\n9    top_n=100,\n10)\n11\n12print(f\"Map ID: {sitemap.id}\")\n13\n14# Iterate all URLs (handles pagination automatically)\n15for url in sitemap.urls():\n16    print(url)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const map = await client.maps.create({\n7  url: 'https://docs.olostep.com',\n8  includeUrls: ['/features/**'],\n9  topN: 100,\n10})\n11\n12console.log(`Map ID: ${map.id}`)\n13\n14// Iterate all URLs (handles pagination automatically)\n15for await (const url of map.urls()) {\n16  console.log(url)\n17}\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/maps\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"url\": \"https://docs.olostep.com\",\n6    \"include_urls\": [\"/features/**\"],\n7    \"top_n\": 100\n8  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/maps)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nProcess up to 10k URLs in one batch. Get results in 5-8 mins\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-5-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-5-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-5-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6batch = client.batches.create(\n7    urls=[\\\n8        {\"custom_id\": \"item-1\", \"url\": \"https://www.google.com/search?q=stripe&gl=us&hl=en\"},\\\n9        {\"custom_id\": \"item-2\", \"url\": \"https://www.google.com/search?q=paddle&gl=us&hl=en\"},\\\n10    ],\n11    parser=\"@olostep/google-search\",\n12)\n13\n14print(batch.id, batch.status)\n15\n16# Wait and iterate results (auto-waits for completion)\n17for item in batch.items():\n18    content = item.retrieve([\"json\"])\n19    print(item.url, item.custom_id)\n20    print(content.json_content)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const batch = await client.batches.create([\\\n7  { url: 'https://www.google.com/search?q=stripe&gl=us&hl=en', customId: 'item-1' },\\\n8  { url: 'https://www.google.com/search?q=paddle&gl=us&hl=en', customId: 'item-2' },\\\n9], {\n10  parser: '@olostep/google-search',\n11})\n12\n13console.log(batch.id, batch.total_urls)\n14\n15// Wait and iterate results (auto-waits for completion)\n16for await (const item of batch.items()) {\n17  const content = await item.retrieve(['json'])\n18  console.log(item.url, item.custom_id)\n19  console.log(content.json_content)\n20}\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/batches\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"items\": [\\\n6      {\"custom_id\": \"item-1\", \"url\": \"https://www.google.com/search?q=stripe&gl=us&hl=en\"},\\\n7      {\"custom_id\": \"item-2\", \"url\": \"https://www.google.com/search?q=paddle&gl=us&hl=en\"}\\\n8    ],\n9    \"parser\": {\"id\": \"@olostep/google-search\"}\n10  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/batches)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nSemantically search the Web\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-6-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-6-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-6-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6search = client.searches.create(\"Latest updates with SpaceX\")\n7\n8print(search.id, len(search.links))\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const search = await client.searches.create('Latest updates with SpaceX')\n7\n8console.log(search.id, search.links.length)\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/searches\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"query\": \"Latest updates with SpaceX\"\n6  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/search)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nGet answers from the Web\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-7-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-7-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-7-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6answer = client.answers.create(\n7    task=\"What does Olostep do and what is its core offering?\",\n8    json_format={\"company\": \"\", \"what_it_does\": \"\", \"core_offering\": \"\"},\n9)\n10\n11print(answer.json_content)\n12print(answer.sources)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const answer = await client.answers.create({\n7  task: 'What does Olostep do and what is its core offering?',\n8  jsonFormat: { company: '', what_it_does: '', core_offering: '' },\n9})\n10\n11console.log(answer.json_content)\n12console.log(answer.sources)\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/answers\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"task\": \"What does Olostep do and what is its core offering?\",\n6    \"json\": {\"company\": \"\", \"what_it_does\": \"\", \"core_offering\": \"\"}\n7  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/answers)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nMonitor pages on a schedule and get change alerts\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-8-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-8-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-8-data-w-pane-2)\n\n```python\n1import requests\n2import json\n3\n4API_KEY = \"<YOUR_API_KEY>\"\n5API_URL = \"https://api.olostep.com/v1\"\n6\n7# Create a monitor\n8payload = {\n9    \"query\": \"Alert me when Tesla stock price is above $500\",\n10    \"frequency\": \"every hour\",\n11    \"email\": \"alerts@example.com\"\n12}\n13\n14headers = {\n15    \"Authorization\": f\"Bearer {API_KEY}\",\n16    \"Content-Type\": \"application/json\"\n17}\n18\n19response = requests.post(f\"{API_URL}/monitors\", headers=headers, json=payload)\n20monitor = response.json()\n21monitor_id = monitor['id']\n22\n23print(f\"Monitor created: {monitor_id}\")\n24print(f\"Status: {monitor['status']}\")\n25\n26# List all monitors\n27monitors = requests.get(f\"{API_URL}/monitors\", headers=headers).json()\n28for m in monitors['monitors']:\n29    print(f\"{m['id']}: {m['url']} ({m['frequency']})\")\n30\n31# Get monitor details\n32details = requests.get(f\"{API_URL}/monitors/{monitor_id}\", headers=headers).json()\n33print(json.dumps(details, indent=2))\n34\n35# Delete a monitor\n36requests.delete(f\"{API_URL}/monitors/{monitor_id}\", headers=headers)\n37print(f\"Monitor {monitor_id} deleted\")\n```\n\n```javascript\n1const API_URL = 'https://api.olostep.com/v1'\n2const headers = {\n3  'Authorization': 'Bearer <YOUR_API_KEY>',\n4  'Content-Type': 'application/json'\n5}\n6\n7// Create a monitor\n8const res = await fetch(`${API_URL}/monitors`, {\n9  method: 'POST',\n10  headers,\n11  body: JSON.stringify({\n12    query: 'Alert me when Tesla stock price is above $500',\n13    frequency: 'every hour',\n14    email: 'alerts@example.com'\n15  })\n16})\n17\n18const monitor = await res.json()\n19console.log(`Monitor created: ${monitor.id}`)\n20console.log(`Status: ${monitor.status}`)\n21\n22// List all monitors\n23const monitors = await fetch(`${API_URL}/monitors`, { headers }).then(r => r.json())\n24monitors.monitors.forEach(m => console.log(`${m.id}: ${m.url} (${m.frequency})`))\n25\n26// Get monitor details\n27const details = await fetch(`${API_URL}/monitors/${monitor.id}`, { headers }).then(r => r.json())\n28console.log(details)\n29\n30// Delete a monitor\n31await fetch(`${API_URL}/monitors/${monitor.id}`, { method: 'DELETE', headers })\n32console.log(`Monitor ${monitor.id} deleted`)\n```\n\n```bash\n1# Create a monitor\n2curl -s -X POST \"https://api.olostep.com/v1/monitors\" \\\n3  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n4  -H \"Content-Type: application/json\" \\\n5  -d '{\n6    \"query\": \"Track changes in product pricing and stock information\",\n7    \"url\": \"https://example.com/products/widget-pro\",\n8    \"frequency\": \"daily\",\n9    \"email\": \"alerts@example.com\"\n10  }'\n11\n12# List all monitors\n13curl -s \"https://api.olostep.com/v1/monitors\" \\\n14  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n15\n16# Get monitor details (replace <MONITOR_ID>)\n17curl -s \"https://api.olostep.com/v1/monitors/<MONITOR_ID>\" \\\n18  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n19\n20# Delete a monitor (replace <MONITOR_ID>)\n21curl -s -X DELETE \"https://api.olostep.com/v1/monitors/<MONITOR_ID>\" \\\n22  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/monitors)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\n## /scrapes\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nTurn any URL into LLM-ready Markdown, HTML, screenshots, PDFs, or structured JSON. Handle JS-rendered pages, actions, and extraction workflows without maintaining browsers, proxies, or brittle scrapers.\n\n## /crawls\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nCrawl websites at scale, collect content from subpages, control depth and URL patterns, and retrieve clean HTML or Markdown for indexing, enrichment, RAG, and AI workflows.\n\n## /maps\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nDiscover every URL on a website using sitemaps and on-page links. Filter by path patterns, paginate large results, and prepare clean URL lists for SEO, crawls, and batches.\n\n## /batches\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nProcess up to 10k concurrent URLs in a single batch in 5-8 mins to get clean web data and aggregate content. Run many batches in parallel to scale to millions of concurrent requests.\n\n## /searches\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nAsk natural-language questions and get AI-generated answers grounded in web sources. Return validated data in the JSON shape you want, with NOT\\_FOUND when facts cannot be verified.\n\n## /answers\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nCreate scheduled web monitors from natural-language instructions. Track for changes on a single page or across the web, deltas, extract structured insights, and get chance notifications by email, webhook, or SMS.\n\n## /monitors\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nTurn recurring website extraction into fast, cost-efficient structured JSON. Use pre-built parsers or create custom parsers for deterministic data pipelines. Use in conjunction with scrapes, crawls, and batches.\n\n[/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-0) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-1) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-2) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-3) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-4) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-5) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-6)\n\n[**Flexible** \\\\\n\\\\\nGet data as HTML, Markdown, text, PDF, JSON, screenshots, or raw bytes.\\\\\n\\\\\n**Structured** \\\\\n\\\\\nExtract clean, structured data using parsers or AI-powered LLM extraction.\\\\\n\\\\\n**Managed** \\\\\n\\\\\nStop managing headless browsers, proxies, and CAPTCHAs. \\\\\n\\\\\nYour Request\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nhttps:/olostep.com/\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nExtract the legal name, mission, and features\\\\\n\\\\\nScraping Infrastructure\\\\\n\\\\\n![](https://www.olostep.com/images/window-check.svg)\\\\\n\\\\\nBrowsers\\\\\n\\\\\n![](https://www.olostep.com/images/reload-window.svg)\\\\\n\\\\\nProxies\\\\\n\\\\\n![](https://www.olostep.com/images/puzzle.svg)\\\\\n\\\\\nCAPTCHAs\\\\\n\\\\\nStructured Result\\\\\n\\\\\n![](https://www.olostep.com/images/page.svg)\\\\\n\\\\\nFormats\\\\\n\\\\\nHTML\\\\\n\\\\\nMarkdown\\\\\n\\\\\nText\\\\\n\\\\\nPDF\\\\\n\\\\\nRaw Bytes\\\\\n\\\\\nScreenshot\\\\\n\\\\\n![](https://www.olostep.com/images/code-brackets-1.svg)\\\\\n\\\\\nStructured Data\\\\\n\\\\\n{ \\\\\n\\\\\n \u00a0\u00a0\u00a0\"name\": 'Olostep Technologies', \\\\\n\\\\\n\u00a0\u00a0\u00a0\u00a0\"mission\": 'Build infrastruct...', \\\\\n\\\\\n\u00a0\u00a0\u00a0\u00a0\"features\": 'monitors, batch...'\\\\\n\\\\\n}](https://www.olostep.com/playground)\n\n[**Scalable** \\\\\n\\\\\nBuilt for large-scale crawling across blogs, documentation, and websites.\\\\\n\\\\\n**Controlled** \\\\\n\\\\\nDefine crawl depth and target only URLs matching specific patterns.\\\\\n\\\\\n**Notified** \\\\\n\\\\\nReceive webhook alerts automatically when crawling jobs complete.\\\\\n\\\\\nhttps://docs.olostep.com/\\\\\n\\\\\nDepth 1\\\\\n\\\\\nDepth 2\\\\\n\\\\\nDepth 3\\\\\n\\\\\nCrawled/ Included\\\\\n\\\\\nExcluded by Rules](https://www.olostep.com/playground?q=crawl)\n\n[**Complete** \\\\\n\\\\\nDiscover all URLs across a website, including pages beyond sitemaps.\\\\\n\\\\\n**Customizable** \\\\\n\\\\\nInclude or exclude paths using flexible URL pattern matching.\\\\\n\\\\\n**Insightful** \\\\\n\\\\\nExplore website structure and identify URLs worth scraping next.\\\\\n\\\\\nhttps://www.olostep.com/\\\\\n\\\\\n![](https://www.olostep.com/images/folder-1.svg)\\\\\n\\\\\n/docs\\\\\n\\\\\n/get-started/welcome\\\\\n\\\\\n/features/maps\\\\\n\\\\\n/integrations/n8n\\\\\n\\\\\n....\\\\\n\\\\\n![](https://www.olostep.com/images/post.svg)\\\\\n\\\\\n/blog\\\\\n\\\\\n/monitors-api\\\\\n\\\\\n/parsers-vs-llm\\\\\n\\\\\n/olostep-orthogonal\\\\\n\\\\\n....\\\\\n\\\\\n![](https://www.olostep.com/images/shop.svg)\\\\\n\\\\\n/store\\\\\n\\\\\n/google-search\\\\\n\\\\\n/brave-search\\\\\n\\\\\n/bing-search\\\\\n\\\\\n....\\\\\n\\\\\n![](https://www.olostep.com/images/code-brackets-square.svg)\\\\\n\\\\\n/api-ref\\\\\n\\\\\n/scrapes/create\\\\\n\\\\\n/batches/create\\\\\n\\\\\n/crawls/create\\\\\n\\\\\n....](https://www.olostep.com/playground?q=map)\n\n[**Concurrent** \\\\\n\\\\\nProcess up to 10k concurrent URLs in a single batch in 5-8 mins.\\\\\n\\\\\n**Massive** \\\\\n\\\\\nRun multiple batches simultaneously to process millions of URLs.\\\\\n\\\\\n**Unique** \\\\\n\\\\\nUnique feature on the market purpose-built for large-scale concurrent processing.\\\\\n\\\\\nInput: URLs\\\\\n\\\\\n1,000,000\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-1\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-2\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-3\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-4\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-5\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-6\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-n\\\\\n\\\\\nConcurrent Batch Processing\\\\\n\\\\\nBatch 1\\\\\n\\\\\nProcessing\\\\\n\\\\\nBatch 2\\\\\n\\\\\nProcessing\\\\\n\\\\\nBatch 3\\\\\n\\\\\nProcessing\\\\\n\\\\\nBatch n\\\\\n\\\\\nProcessing\\\\\n\\\\\n![](https://www.olostep.com/images/check-circle.svg)\\\\\n\\\\\nURLs Processed\\\\\n\\\\\n1,000,000\\\\\n\\\\\n![](https://www.olostep.com/images/apple-shortcuts-2.svg)\\\\\n\\\\\nConcurrent Batches\\\\\n\\\\\n128\\\\\n\\\\\n![](https://www.olostep.com/images/dashboard-speed.svg)\\\\\n\\\\\nThroughput\\\\\n\\\\\n48,752 /min\\\\\n\\\\\n![](https://www.olostep.com/images/clock.svg)\\\\\n\\\\\nTotal time\\\\\n\\\\\n20m 14s](https://www.olostep.com/playground?q=batch)\n\n[**Natural** \\\\\n\\\\\nGet relevant links with titles and descriptions using natural language queries.\\\\\n\\\\\n**Expandable** \\\\\n\\\\\nCombine search and scraping into a single API call for faster data retrieval.\\\\\n\\\\\n**Focused** \\\\\n\\\\\nInclude or exclude domains to refine search results precisely.\\\\\n\\\\\nYour Query\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nBest AI Agents Framework 2026\\\\\n\\\\\nInclude\\\\\n\\\\\nLangGraph\\\\\n\\\\\nCrewAi\\\\\n\\\\\nExclude\\\\\n\\\\\nOpenAI SDK\\\\\n\\\\\nAutoGen\\\\\n\\\\\n![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nLangGraph\\\\\n\\\\\n![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nAutoGen\\\\\n\\\\\n![](https://www.olostep.com/images/170677839.png)\\\\\n\\\\\nCrewAI\\\\\n\\\\\n![](https://www.olostep.com/images/2.jpg)\\\\\n\\\\\nOpenAI SDK\\\\\n\\\\\nStructured Results\\\\\n\\\\\n![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nhttps:/langchain...\\\\\n\\\\\nLangGraph\\\\\n\\\\\nBalance agent control with agency...\\\\\n\\\\\n![](https://www.olostep.com/images/170677839.png)\\\\\n\\\\\nhttps://crewai.com/\\\\\n\\\\\nCrewAI\\\\\n\\\\\nThe open platform that accelerates agent...](https://www.olostep.com/playground?q=search)\n\n[**Grounded** \\\\\n\\\\\nSearch the web and generate answers from real-world sources.\\\\\n\\\\\n**Structured** \\\\\n\\\\\nReturn validated results and sources in the exact JSON shape you need.\\\\\n\\\\\n**Enrich** \\\\\n\\\\\nEnhance products, datasets, and spreadsheets with web data.\\\\\n\\\\\nYour Question\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nFind YC startups building voice AI\\\\\n\\\\\nProcessing Layer\\\\\n\\\\\n![](https://www.olostep.com/images/search_1.svg)\\\\\n\\\\\nSearch\\\\\n\\\\\n![](https://www.olostep.com/images/spark.svg)\\\\\n\\\\\nClean\\\\\n\\\\\n![](https://www.olostep.com/images/check-circle.svg)\\\\\n\\\\\nValidate\\\\\n\\\\\nStructured Results\\\\\n\\\\\n{\"results\": \\[\\\\\n\\\\\n \u00a0 \u00a0{\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"company\": \"Retell AI\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"yc\\_batch\": \"W24\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"category\": \"Voice Agents\"\\\\\n\\\\\n \u00a0 \u00a0},\\\\\n\\\\\n \u00a0 \u00a0{\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"company\": \"Vapi\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"yc\\_batch\": \"W21\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"category\": \"Voice Infrastructure\"\\\\\n\\\\\n \u00a0 \u00a0}\\\\\n\\]}](https://www.olostep.com/playground?q=answer)\n\n[**Scheduled** \\\\\n\\\\\nTrack website changes automatically on a recurring schedule.\\\\\n\\\\\n**Alerts** \\\\\n\\\\\nGet notified via email, SMS, webhooks, or custom channels.\\\\\n\\\\\n**Deterministic** \\\\\n\\\\\nCreate monitors with natural language and run them deterministically.\\\\\n\\\\\nSetup to monitor\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nMonitor the OpenAI pricing page & alert me if any pricing changes\\\\\n\\\\\nSchedule\\\\\n\\\\\nEvery 6 hours\\\\\n\\\\\nAlert Delivered\\\\\n\\\\\nChange Summary\\\\\n\\\\\n$20 -> $22 (+$2)\\\\\n\\\\\n![](https://www.olostep.com/images/mail_1.svg)\\\\\n\\\\\nEmail\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nWebhook\\\\\n\\\\\n![](https://www.olostep.com/images/chat-bubble.svg)\\\\\n\\\\\nSms](https://www.olostep.com/dashboard/monitors)\n\n## Data tailored to your industry\n\nSee how Olostep powers AI platforms, sales lead enrichment, deep research, competitive intelligence, and SEO teams with one API.\n\n### Deep Search\n\nAccess custom, hyper-specialized B2B indexes for your industry to search and extract comprehensive data beyond what general web indexes cover\n\nAdd Enrichment\n\n![](https://www.olostep.com/images/group.svg)\n\nSearch for persons at company\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/mail.svg)\n\nFind business emails\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/coins.svg)\n\nFind annual revenues\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/suitcase.svg)\n\nSearch job openings\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/map-pin.svg)\n\nFind address\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Recruiting\n\nIdentify, research, and validate candidates faster with intelligence and data aggregated from top-quality profiles and specialist web sources.\n\nRecruiting Pipeline\n\n![](https://www.olostep.com/images/search-engine.svg)\n\nSourcing\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/filter.svg)\n\nPreprocessing\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/label.svg)\n\nLabeling\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/stats-down-square.svg)\n\nEvaluation\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/rocket.svg)\n\nDeployment\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Power AI applications\n\nGet clean, structured data from any website as markdown, html, screenshot, etc. to power your AI application and workflows\n\nExtract Content\n\n![](https://www.olostep.com/images/html5.svg)\n\nExtract as Markdown / HTML\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/code-brackets.svg)\n\nExtract as JSON\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/code.svg)\n\nExtract Code\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/multiple-pages.svg)\n\nExtract as Text / PDF\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/screenshot.svg)\n\nExtract as Screenshot\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Monitor the Web\n\nMonitor any webpage for DOM changes, stock availability, price changes, job openings or fresh content. Run automatically on a schedule and get alerted\n\nMonitor Webpage\n\n![](https://www.olostep.com/images/candlestick-chart.svg)\n\nMonitor pricing & stock\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/eye.svg)\n\nMonitor content & DOM changes\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/suitcase.svg)\n\nMonitor job openings\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/leaderboard-star.svg)\n\nMonitor reviews & ratings\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/alarm.svg)\n\nSet schedule & alerts\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Automate data pipelines\n\nAutomate complex data pipelines with the /agents endpoint through natural language prompts. You can also pass your own internal knowledge as context\n\nNatural language to data pipelines\n\n![](https://www.olostep.com/images/search-engine.svg)\n\nPower your search product\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/eye.svg)\n\nTrack portfolio companies\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/coins.svg)\n\nBuild market intelligence\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/stats-down-square.svg)\n\nResearch companies at scale\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/suitcase.svg)\n\nAutomate GTM and recruiting\n\n![](https://www.olostep.com/images/Group-2.svg)\n\nAutomate data pipelines\n\nAutomate complex data pipelines with the /agents endpoint through natural language prompts. You can also pass your own internal knowledge as context\n\n![](https://www.olostep.com/images/image-5.png)\n\n[![](https://www.olostep.com/images/nav-arrow-left.svg)](https://www.olostep.com/#)[![](https://www.olostep.com/images/nav-arrow-right.svg)](https://www.olostep.com/#)\n\n## Pricing that Makes Sense\n\nMost cost-effective web data API on the market\n\n**No credit card required**\n\n### Trial\n\n$0\n\nIncludes:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n500 successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\nAll requests are JS rendered + utilizing residential IP addresses\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\nLow rate limits\n\n[Get started](https://www.olostep.com/dashboard/plans)\n\nCOST/1K$1.800\n\n### Starter\n\n$9\n\n/ month\n\nEverything in Free, Plus:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n5000 successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n150 concurrent requests\n\n[Purchase now](https://www.olostep.com/dashboard/plans)\n\nCOST/1K**$0.495**\n\n### Standard\n\n$99\n\n/ month\n\nEverything in Starter, Plus:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n200K successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n500 concurrent requests\n\n[Purchase now](https://www.olostep.com/dashboard/plans)\n\nCOST/1K$ **0.399**\n\n### Scale\n\n$399\n\n/ month\n\nEverything in Standard, Plus:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n1 Million successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\nAI-powered Browser Automations\n\n[Purchase now](https://www.olostep.com/dashboard/plans)\n\n## Top-ups\n\nHave spiky usage or don't like subscriptions?\n\nYou can buy credit packs. They are valid for 6 months.\n\nCredit pack\n\n### 10k credits\n\n**$20**\n\n[Purchase Credit Pack](https://www.olostep.com/dashboard/plans)\n\nCredit pack\n\n### 250k credits\n\n**$200**\n\n[Purchase Credit Pack](https://www.olostep.com/dashboard/plans)\n\nCredit pack\n\n### 2M credits\n\n**$1000**\n\n[Purchase Credit Pack](https://www.olostep.com/dashboard/plans)\n\n### Enterprise\n\nHundreds of millions of credits with enterprise-grade reliability. We offer custom discounts\n\n[Contact Sales](https://www.olostep.com/contact-sales)\n\n## Trusted by Amazing Teams Building the Future of AI\n\n![](https://www.olostep.com/images/michelle-pic.jpg)\n\nMichelle Julia\n\nCo-founder & CEO Aurium\n\nOlostep is the best!!! We automated entire data pipelines with just a prompt\n\n![](https://www.olostep.com/images/richard-he.avif)\n\nRichard He\n\nCo-founder & CEO Openmart\n\nOlostep has become the default Web Layer infrastructure for our company\n\n![](https://www.olostep.com/images/mx-pic.jpeg)\n\nMax Brodeur-Urbas\n\nCo-founder & CEO Gumloop\n\nOlostep works like a charm! And your customer service is exceptional\n\n![](https://www.olostep.com/images/rob-pic.jpeg)\n\nRob Hayes\n\nCo-founder Merchkit\n\nOlostep lets us turn any website into an API. Great product, great people\n\n![](https://www.olostep.com/images/brandon-civilgrid.webp)\n\nBrandon Cohen\n\nCo-founder & CTO CivilGrid\n\nI highly recommend Olostep, great product!\n\n![](https://www.olostep.com/images/michelle-pic.jpg)\n\nMichelle Julia\n\nCo-founder & CEO Aurium\n\nOlostep is the best!!! We automated entire data pipelines with just a prompt\n\n![](https://www.olostep.com/images/richard-he.avif)\n\nRichard He\n\nCo-founder & CEO Openmart\n\nOlostep has become the default Web Layer infrastructure for our company\n\n![](https://www.olostep.com/images/mx-pic.jpeg)\n\nMax Brodeur-Urbas\n\nCo-founder & CEO Gumloop\n\nOlostep works like a charm! And your customer service is exceptional\n\n![](https://www.olostep.com/images/rob-pic.jpeg)\n\nRob Hayes\n\nCo-founder Merchkit\n\nOlostep lets us turn any website into an API. Great product, great people\n\n![](https://www.olostep.com/images/brandon-civilgrid.webp)\n\nBrandon Cohen\n\nCo-founder & CTO CivilGrid\n\nI highly recommend Olostep, great product!\n\n![](https://www.olostep.com/images/berk-pic.jpeg)\n\n[Berk Serbetcioglu](https://www.berkserbetcioglu.com/)\n\nCo-founder & CEO Gedd.it\n\nWe verify coupon codes at scale. Love Olostep. It works on any e-commerce\n\n![](https://www.olostep.com/images/trevor-west.jpeg)\n\nTrevor West\n\nCo-founder & CEO Podqi\n\nOlostep is the best API to search, extract, and structure data from the Web. Happy to be customers\n\n![](https://www.olostep.com/images/rida_pic.jpg)\n\nRida Naveed\n\nCo-founder Zecento\n\nWe use /batches combined with parsers and it's magical how we can get structured data at large scale\n\n![](https://www.olostep.com/images/kieran-plots.jpeg)\n\nKieran V.\n\nGrowth PlotsEvents\n\nOlostep allowed us to search and structure events data across the Web\n\n![](https://www.olostep.com/images/paul-mit.jpg)\n\nPaul Mit\n\nFounder Foundbase\n\nReliable and cost-effective API for working with data. Congrats on the cool product\n\n![](https://www.olostep.com/images/berk-pic.jpeg)\n\n[Berk Serbetcioglu](https://www.berkserbetcioglu.com/)\n\nCo-founder & CEO Gedd.it\n\nWe verify coupon codes at scale. Love Olostep. It works on any e-commerce\n\n![](https://www.olostep.com/images/trevor-west.jpeg)\n\nTrevor West\n\nCo-founder & CEO Podqi\n\nOlostep is the best API to search, extract, and structure data from the Web. Happy to be customers\n\n![](https://www.olostep.com/images/rida_pic.jpg)\n\nRida Naveed\n\nCo-founder Zecento\n\nWe use /batches combined with parsers and it's magical how we can get structured data at large scale\n\n![](https://www.olostep.com/images/kieran-plots.jpeg)\n\nKieran V.\n\nGrowth PlotsEvents\n\nOlostep allowed us to search and structure events data across the Web\n\n![](https://www.olostep.com/images/paul-mit.jpg)\n\nPaul Mit\n\nFounder Foundbase\n\nReliable and cost-effective API for working with data. Congrats on the cool product\n\n## Connect Olostep to Your AI Stack\n\nOfficial Olostep integrations. Add web scraping,\n\ncrawling and AI-powered search to any tool in your stack.\n\n[![](https://www.olostep.com/images/cursor.webp)\\\\\n\\\\\nCursor](https://www.olostep.com/#) [![](https://www.olostep.com/images/claude.jpeg)\\\\\n\\\\\nClaude](https://www.olostep.com/#) [![](https://www.olostep.com/images/n8n.webp)\\\\\n\\\\\nn8n](https://www.olostep.com/#) [![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nLangChain](https://www.olostep.com/#) [![](https://www.olostep.com/images/4.jpg)\\\\\n\\\\\nApify](https://www.olostep.com/#) [![](https://www.olostep.com/images/149120496.png)\\\\\n\\\\\nMastra](https://www.olostep.com/#) [![](https://www.olostep.com/images/zapier.webp)\\\\\n\\\\\nZapier](https://www.olostep.com/#) [![](https://www.olostep.com/images/relay.webp)\\\\\n\\\\\nRelay](https://www.olostep.com/#) [![](https://www.olostep.com/images/windsurf.webp)\\\\\n\\\\\nWindsurf](https://www.olostep.com/#) [![](https://www.olostep.com/images/3.jpg)\\\\\n\\\\\nCline](https://www.olostep.com/#) [![](https://www.olostep.com/images/vscode.webp)\\\\\n\\\\\nVS Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/1.jpg)\\\\\n\\\\\nGemini CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/2.jpg)\\\\\n\\\\\nOpenAI Codex](https://www.olostep.com/#) [![](https://www.olostep.com/images/cursor.webp)\\\\\n\\\\\nCursor](https://www.olostep.com/#) [![](https://www.olostep.com/images/claude.jpeg)\\\\\n\\\\\nClaude](https://www.olostep.com/#) [![](https://www.olostep.com/images/n8n.webp)\\\\\n\\\\\nn8n](https://www.olostep.com/#) [![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nLangChain](https://www.olostep.com/#) [![](https://www.olostep.com/images/4.jpg)\\\\\n\\\\\nApify](https://www.olostep.com/#) [![](https://www.olostep.com/images/149120496.png)\\\\\n\\\\\nMastra](https://www.olostep.com/#) [![](https://www.olostep.com/images/zapier.webp)\\\\\n\\\\\nZapier](https://www.olostep.com/#) [![](https://www.olostep.com/images/relay.webp)\\\\\n\\\\\nRelay](https://www.olostep.com/#) [![](https://www.olostep.com/images/windsurf.webp)\\\\\n\\\\\nWindsurf](https://www.olostep.com/#) [![](https://www.olostep.com/images/3.jpg)\\\\\n\\\\\nCline](https://www.olostep.com/#) [![](https://www.olostep.com/images/vscode.webp)\\\\\n\\\\\nVS Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/1.jpg)\\\\\n\\\\\nGemini CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/2.jpg)\\\\\n\\\\\nOpenAI Codex](https://www.olostep.com/#)\n\n[![](https://www.olostep.com/images/qwen.jpg)\\\\\n\\\\\nQwen](https://www.olostep.com/#) [![](https://www.olostep.com/images/smithery.jpg)\\\\\n\\\\\nSmithery](https://www.olostep.com/#) [![](https://www.olostep.com/images/amazonaws.webp)\\\\\n\\\\\nAmazon Q](https://www.olostep.com/#) [![](https://www.olostep.com/images/ampcode.webp)\\\\\n\\\\\nAmp](https://www.olostep.com/#) [![](https://www.olostep.com/images/augment.jpg)\\\\\n\\\\\nAugment](https://www.olostep.com/#) [![](https://www.olostep.com/images/boltai.webp)\\\\\n\\\\\nBoltAI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Bun.jpg)\\\\\n\\\\\nBun](https://www.olostep.com/#) [![](https://www.olostep.com/images/deno-1.webp)\\\\\n\\\\\nDeno](https://www.olostep.com/#) [![](https://www.olostep.com/images/copilot.webp)\\\\\n\\\\\nCopilot](https://www.olostep.com/#) [![](https://www.olostep.com/images/124303983.png)\\\\\n\\\\\nCrush](https://www.olostep.com/#) [![](https://www.olostep.com/images/docker.webp)\\\\\n\\\\\nDocker](https://www.olostep.com/#) [![](https://www.olostep.com/images/jetbrains.webp)\\\\\n\\\\\nJetBrains](https://www.olostep.com/#) [![](https://www.olostep.com/images/kiro.webp)\\\\\n\\\\\nKiro](https://www.olostep.com/#) [![](https://www.olostep.com/images/qwen.jpg)\\\\\n\\\\\nQwen](https://www.olostep.com/#) [![](https://www.olostep.com/images/smithery.jpg)\\\\\n\\\\\nSmithery](https://www.olostep.com/#) [![](https://www.olostep.com/images/amazonaws.webp)\\\\\n\\\\\nAmazon Q](https://www.olostep.com/#) [![](https://www.olostep.com/images/ampcode.webp)\\\\\n\\\\\nAmp](https://www.olostep.com/#) [![](https://www.olostep.com/images/augment.jpg)\\\\\n\\\\\nAugment](https://www.olostep.com/#) [![](https://www.olostep.com/images/boltai.webp)\\\\\n\\\\\nBoltAI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Bun.jpg)\\\\\n\\\\\nBun](https://www.olostep.com/#) [![](https://www.olostep.com/images/deno-1.webp)\\\\\n\\\\\nDeno](https://www.olostep.com/#) [![](https://www.olostep.com/images/copilot.webp)\\\\\n\\\\\nCopilot](https://www.olostep.com/#) [![](https://www.olostep.com/images/124303983.png)\\\\\n\\\\\nCrush](https://www.olostep.com/#) [![](https://www.olostep.com/images/docker.webp)\\\\\n\\\\\nDocker](https://www.olostep.com/#) [![](https://www.olostep.com/images/jetbrains.webp)\\\\\n\\\\\nJetBrains](https://www.olostep.com/#) [![](https://www.olostep.com/images/kiro.webp)\\\\\n\\\\\nKiro](https://www.olostep.com/#)\n\n[![](https://www.olostep.com/images/LM.jpg)\\\\\n\\\\\nLM Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/opencode.webp)\\\\\n\\\\\nOpencode](https://www.olostep.com/#) [![](https://www.olostep.com/images/perplexity.webp)\\\\\n\\\\\nPerplexity](https://www.olostep.com/#) [![](https://www.olostep.com/images/qodo.webp)\\\\\n\\\\\nQodo Gen](https://www.olostep.com/#) [![](https://www.olostep.com/images/Roo.jpg)\\\\\n\\\\\nRoo Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/atlassian.webp)\\\\\n\\\\\nRovo Dev CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Traelogo.png)\\\\\n\\\\\nTrae](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nVisual Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/warp.webp)\\\\\n\\\\\nWarp](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nWindows](https://www.olostep.com/#) [![](https://www.olostep.com/images/zed.webp)\\\\\n\\\\\nZed](https://www.olostep.com/#) [![](https://www.olostep.com/images/Zencoder.jpg)\\\\\n\\\\\nZencoder](https://www.olostep.com/#) [![](https://www.olostep.com/images/modelcontextprotocol.webp)\\\\\n\\\\\nMCP client](https://www.olostep.com/#) [![](https://www.olostep.com/images/LM.jpg)\\\\\n\\\\\nLM Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/opencode.webp)\\\\\n\\\\\nOpencode](https://www.olostep.com/#) [![](https://www.olostep.com/images/perplexity.webp)\\\\\n\\\\\nPerplexity](https://www.olostep.com/#) [![](https://www.olostep.com/images/qodo.webp)\\\\\n\\\\\nQodo Gen](https://www.olostep.com/#) [![](https://www.olostep.com/images/Roo.jpg)\\\\\n\\\\\nRoo Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/atlassian.webp)\\\\\n\\\\\nRovo Dev CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Traelogo.png)\\\\\n\\\\\nTrae](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nVisual Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/warp.webp)\\\\\n\\\\\nWarp](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nWindows](https://www.olostep.com/#) [![](https://www.olostep.com/images/zed.webp)\\\\\n\\\\\nZed](https://www.olostep.com/#) [![](https://www.olostep.com/images/Zencoder.jpg)\\\\\n\\\\\nZencoder](https://www.olostep.com/#) [![](https://www.olostep.com/images/modelcontextprotocol.webp)\\\\\n\\\\\nMCP client](https://www.olostep.com/#)\n\n![](https://www.olostep.com/images/flash.svg)\n\n### Connect with your AI agents\n\nA deterministic, repeatable, controllable pipeline that automates any web research workflow and pipeline exactly as you described it.\n\n![](https://www.olostep.com/images/olostep-logo-cropped.svg)\n\n![](https://www.olostep.com/images/n8n.webp)\n\n![](https://www.olostep.com/images/7NLTL1CG1JNK-180x180.jpeg)\n\n![](https://www.olostep.com/images/claude.jpeg)\n\n![](https://www.olostep.com/images/windsurf.webp)\n\n![](https://www.olostep.com/images/perplexity.webp)\n\n![](https://www.olostep.com/images/vscode.webp)\n\n![](https://www.olostep.com/images/terminal.svg)\n\n### Use the Olostep CLI\n\nMap, scrape, crawl, batch process, and generate answers directly from your terminal, with clean JSON output built for scripts, CI pipelines and AI agents.\n\n`npx -y olostep-cli@latest --help`\n\n![](https://www.olostep.com/images/ev-plug.svg)\n\n### Add Olostep to your MCP client\n\nWorks with any product that implements the Model Context Protocol: register Olostep once and call web tools from chat, agents, or IDEs that support MCP.\n\n`{\n\"mcpServers\": {\n \u00a0 \u00a0\"olostep-web\": {\n \u00a0 \u00a0 \u00a0\"command\": \"``npx``\",\n \u00a0 \u00a0 \u00a0\"args\": [\"``-y``\", \"``olostep-mcp``\"],\n \u00a0 \u00a0 \u00a0\"env\": {\n \u00a0 \u00a0 \u00a0 \u00a0\"OLOSTEP_API_KEY\": \"``YOUR_API_KEY``\"\n \u00a0 \u00a0 \u00a0}\n \u00a0 \u00a0}\n}\n}`\n\n## Ready to start?\n\nGet **clean data** for your AI from any website with Olostep\n\nMost cost-effective API. Built for scale\n\n[Start for free](https://www.olostep.com/auth) [See Pricing](https://www.olostep.com/pricing)\n\n[Are you an AI Agent? Get started here](https://www.olostep.com/agent-onboarding/SKILL.md)\n\n## Frequently asked questions\n\nProduct & Capabilities\n\nUsage & Automation\n\nPricing & Plans\n\n### Does Olostep offer a free trial?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYes, Olostep includes a free plan with 500 requests to help you test the API before upgrading. Paid plans start from $9/month and include 5,000 credits per month.\n\nThis gives teams a low-risk way to evaluate Olostep's reliability, scalability and cost-effectiveness before moving to higher-volume usage.\n\n### Can I switch plans after signing up?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYes, you can switch plans at any time. Plans are pro-rated, meaning any unused value from your current plan is carried over to your new plan.\n\nThis ensures you don't pay twice for usage you've already covered, giving you flexibility as your needs grow.\n\n### Can I ask for a refund if I don't use it?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYes. If you're not satisfied with the Olostep API or it doesn't end up being useful for your use case, you can email [info@olostep.com](mailto:info@olostep.com?subject=I%27d%20like%20a%20refund) to request a refund.\n\nIf you cancel after a period of non-use, Olostep can also refund the unused portion of your plan where applicable.\n\n### How can I pay?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYou can pay through Stripe Payment Links. To access billing, you must be logged in to your Olostep account. Go to your dashboard, navigate to Billing & Invoices, click Manage on Stripe, and add your card details. Once your card is added, you\u2019re good to go.\n\n[![](https://www.olostep.com/images/olostep-logo-cropped.svg)](https://www.olostep.com/)\n\nInfrastructure for the Web's second user\n\n[![Olostep - Turn the Web into Clean Data for AI | Product Hunt](https://api.producthunt.com/widgets/embed-image/v1/featured.svg?post_id=1199418&theme=light&t=1788106286872)](https://www.producthunt.com/products/olostep?embed=true&utm_source=badge-featured&utm_medium=badge&utm_campaign=badge-olostep-2)\n\n[![](https://www.olostep.com/images/slack-svgrepo-com.svg)](https://olostep-users.slack.com/join/shared_invite/zt-2bfddyi8h-JzfjOgavg~98DJ1om1B5Lg#/shared-invite/email)\n\n[All services are online](https://status.olostep.com/)\n\nProduct\n\n[Playground](https://www.olostep.com/playground) [Agent](https://www.olostep.com/agent) [Orbit](https://www.olostep.com/orbit) [MCP Server](https://www.olostep.com/integrations/mcp-server) [Integrations](https://www.olostep.com/integrations) [Tools](https://www.olostep.com/tools) [Pricing](https://www.olostep.com/pricing)\n\nDevelopers\n\n[API Reference](https://docs.olostep.com/api-reference/scrapes/create) [API Endpoints](https://www.olostep.com/api-endpoints) [Scrape API](https://www.olostep.com/api-endpoints/scrapes) [Crawl API](https://www.olostep.com/api-endpoints/crawls) [Map API](https://www.olostep.com/api-endpoints/maps) [Batch API](https://www.olostep.com/api-endpoints/batches) [Search API](https://www.olostep.com/api-endpoints/searches) [Answer API](https://www.olostep.com/api-endpoints/answers) [Monitor API](https://www.olostep.com/api-endpoints/monitors) [People and Company Search API](https://www.olostep.com/company-people-search-api)\n\nSolutions\n\n[Tools](https://www.olostep.com/tools) [Connect to Slack](https://www.olostep.com/slack) [Use Cases](https://www.olostep.com/use-cases) [AI Platforms](https://www.olostep.com/use-cases/power-ai-platforms) [Deep Research](https://www.olostep.com/use-cases/deep-research) [Competitive Intelligence](https://www.olostep.com/use-cases/competitive-intelligence) [Sales Lead Enrichment](https://www.olostep.com/use-cases/sales-lead-enrichment) [SEO & GEO Teams](https://www.olostep.com/seo-ai-visibility) [Brand Protection](https://www.olostep.com/brand-protection)\n\nResources\n\n[Documentation](https://docs.olostep.com/get-started/welcome) [Blog](https://www.olostep.com/blog) [Glossary](https://www.olostep.com/glossary) [Changelog](https://www.olostep.com/changelog)\n\nCompany\n\n[About](https://www.olostep.com/about) [Careers](https://www.olostep.com/careers) [Partners](https://www.olostep.com/partners) [Contact Sales](https://www.olostep.com/contact-sales) [AI\u00a0Instructions](https://www.olostep.com/ai-instructions)\n\nLegal\n\nSocials\n\n[X](https://x.com/olostep) [LinkedIn](https://www.linkedin.com/company/olostep/) [Join Slack](https://olostep-users.slack.com/join/shared_invite/zt-2bfddyi8h-JzfjOgavg~98DJ1om1B5Lg#/shared-invite/email)\n\n![](https://www.olostep.com/images/headset-help.svg)\n\nNeed a hand with anything? [Join our Slack](https://olostep-users.slack.com/join/shared_invite/zt-2bfddyi8h-JzfjOgavg~98DJ1om1B5Lg#/shared-invite/email) or send an email to [info@olostep.com](mailto:info@olostep.com?subject=Customer%20support%20request)\n\n\u00a9 2026 Olostep Technologies, Inc.\n\nMade in Italy \\| United States",
      "content_chars": 54156,
      "published_date": null
    },
    {
      "rank": 2,
      "url": "https://docs.olostep.com/integrations/n8n",
      "title": "Olostep + n8n",
      "content": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/integrations/n8n#content-area)\n\nThe verified [Olostep Web Scraper node](https://n8n.io/integrations/olostep-web-scraper/) gives you six operations inside n8n\u2019s visual builder: scrape a URL, search the web, get AI answers, batch-scrape thousands of URLs, crawl a site, or map all its links.[View on n8n \u2192](https://n8n.io/integrations/olostep-web-scraper/)\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#before-you-start)  Before you start\n\n- **An Olostep account with an API key:** [get one free](https://olostep.com/dashboard), no credit card required. Your first 500 credits are included.\n- **n8n running:** either [n8n Cloud](https://n8n.io/cloud/) or a self-hosted instance. Community nodes must be enabled (they are by default on most setups).\n- **No coding required:** everything in this guide is done through n8n\u2019s visual editor.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#setup)  Setup\n\n1\n\nSearch for the Olostep node\n\nOpen any workflow, click **+**, and search for **Olostep**. Select **Olostep Web Scraper** from the results.![Search for Olostep in the n8n node picker](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-1.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=dee8769fcd2a85836d8fb16c12f577ba)\n\n2\n\nInstall the node\n\nClick the result to open the Node details panel, then click **Install node**. n8n will install `n8n-nodes-olostep` and prompt you to restart. Do that before continuing.![Olostep Web Scraper node details with Install node button](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-2.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=2721be5e49905b7428c6a082ed73dd18)\n\nIf **Community Nodes** is disabled for your workspace, an admin needs to enable it first. See the [n8n community nodes guide](https://docs.n8n.io/integrations/community-nodes/installation/).\n\n3\n\nAdd your API key\n\nOpen the Olostep node in your workflow, click **Set up Credential** (in the Parameters tab), add your API key, and click **Save**.![Olostep credentials form in n8n with API Key field](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-3.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=ce6d9f0d70ee4f4a8f167bf1909b7c8d)Get your key from the [Olostep dashboard \u2192](https://olostep.com/dashboard)\n\n4\n\nWire it up and run\n\nConnect the Olostep node to a trigger and any downstream steps, then execute your workflow.![n8n workflow canvas with Schedule Trigger connected to Olostep node](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-4.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=bf0a711ca890d85ed1b8915c50035697)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#actions)  Actions\n\n## Scrape Website\n\nPull content from any URL as Markdown, HTML, JSON, or plain text. Handles JS-rendered pages with optional wait times and country targeting.\n\n## Search\n\nRun a web search and get structured results (titles, URLs, and snippets) as JSON.\n\n## Answers (AI)\n\nAsk a natural-language question and get an answer with cited sources. Useful before LLM nodes when you need grounded responses.\n\n## Batch Scrape URLs\n\nSubmit up to 10,000 URLs in one job, processed in parallel. Returns a `batch_id`; retrieve results asynchronously.\n\n## Create Crawl\n\nStart from a URL, follow links, and scrape all subpages. Good for docs sites, blogs, or full-site ingestion. Returns a `crawl_id`.\n\n## Create Map\n\nGet every URL on a site without scraping content. Use it for discovery before a batch job. Returns a `map_id`.\n\n**Batch, Crawl, and Map are async.** Store the returned ID and use a Wait node or a second workflow to retrieve results once processing completes.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#example-workflow-lead-enrichment-from-google-sheets)  Example workflow: Lead enrichment from Google Sheets\n\n**What it does:** When you paste a company URL into a Google Sheet, this workflow automatically scrapes the company\u2019s website, extracts key information with an AI node, and writes the results back to the same row, turning a blank spreadsheet into a filled-out lead database.**Nodes used:** Google Sheets trigger \u2192 Olostep Scrape Website \u2192 OpenAI \u2192 Code \u2192 Google Sheets update![Lead enrichment workflow in n8n: Google Sheets trigger connected to Olostep, OpenAI, Code, and Google Sheets update nodes](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/workflow-example.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=60e725c70903c2b3cf6918f2ca486e70)\n\n* * *\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-1-set-up-your-google-sheet)  Step 1: Set up your Google Sheet\n\nCreate a sheet with these columns: `Company URL`, `Industry`, `Description`, `Company Size`, `Enriched`. The workflow reads from `Company URL` and fills in the rest.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-2-add-a-google-sheets-trigger)  Step 2: Add a Google Sheets trigger\n\nIn n8n, add a **Google Sheets** trigger node. Set the event to **Row Added**, point it at your sheet, and set it to watch the `Company URL` column. Now every time you paste a new URL into the sheet, this workflow fires.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-3-add-olostep-scrape-website)  Step 3: Add Olostep Scrape Website\n\nConnect an **Olostep Web Scraper** node after the trigger. Set:\n\n- **Action:** Scrape Website\n- **URL:**`{{ $json[\"Company URL\"] }}` (pulls the URL from the new row)\n- **Output Format:** Markdown\n\nMarkdown works best here because it strips navigation, ads, and boilerplate. The AI node in the next step gets clean prose about the company instead of raw HTML noise.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-4-add-an-openai-node)  Step 4: Add an OpenAI node\n\nConnect an **OpenAI** node. Set the model to `gpt-4o-mini` (fast and cheap for extraction tasks) and use this prompt:\n\n```\nYou are a sales researcher. Based on the company website content below, extract:\n1. Industry (one phrase, e.g. \"B2B SaaS\", \"E-commerce\", \"Healthcare\")\n2. One-sentence company description (max 20 words)\n3. Estimated company size (Startup / SMB / Mid-market / Enterprise)\n\nReturn only a JSON object with keys: industry, description, company_size.\n\nWebsite content:\n{{ $json.markdownContent }}\n```\n\nThe `markdownContent` field is what Olostep returns from the scrape, as clean plain text.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-5-parse-the-ai-response-and-write-back)  Step 5: Parse the AI response and write back\n\nAdd a **Code** node to parse the JSON from OpenAI:\n\n```\nconst parsed = JSON.parse($input.first().json.message.content);\nreturn [{ json: parsed }];\n```\n\nThen connect a **Google Sheets** node set to **Update Row**. Map the columns:\n\n- `Industry` \u2192 `{{ $json.industry }}`\n- `Description` \u2192 `{{ $json.description }}`\n- `Company Size` \u2192 `{{ $json.company_size }}`\n- `Enriched` \u2192 `Yes`\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#what-you-get)  What you get\n\nPaste a URL like `https://notion.so` into your sheet, and within ~10 seconds the row fills in:\n\n| Company URL | Industry | Description | Company Size | Enriched |\n| --- | --- | --- | --- | --- |\n| [https://notion.so](https://notion.so/) | Productivity SaaS | All-in-one workspace for notes, docs, and databases | Mid-market | Yes |\n\nFrom here you can extend this workflow: add a Slack notification when enrichment completes, filter by industry before writing back, or replace Google Sheets with HubSpot to update contacts directly.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#templates)  Templates\n\nReady-to-import n8n workflows built with Olostep:\n\n[**Crawl docs \u2192 AI knowledge base** \\\\\n\\\\\nCrawl documentation sites with Olostep and structure the output into an AI-ready knowledge base.](https://www.n8n.io/workflows/13436-crawl-documentation-sites-and-build-an-ai-knowledge-base-with-olostep/)\n\n[**Google Maps leads \u2192 decision-maker enrichment** \\\\\n\\\\\nScrape business leads from Google Maps and enrich them with decision-maker details.](https://n8n.io/workflows/11086-scrape-business-leads-from-google-maps-and-extract-decision-maker-info-with-olostep/)\n\n[**Mine user complaints \u2192 insight report** \\\\\n\\\\\nAnalyze complaints with Olostep + Gemini and generate structured insight reports in Google Docs.](https://www.n8n.io/workflows/13435-mine-user-complaints-and-generate-insight-reports-with-olostep-gemini-and-google-docs/)\n\n[**Amazon product extraction \u2192 Google Sheets** \\\\\n\\\\\nExtract Amazon product URLs and metadata with Olostep, then sync the results to Sheets.](https://n8n.io/workflows/11158-extract-amazon-product-data-to-sheets-with-olostep-api/)\n\n[Browse all Olostep workflows on n8n.io \u2192](https://n8n.io/workflows/?q=olostep)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#parsers)  Parsers\n\nAdd a parser ID to the **Parser** field on any Scrape or Batch action to get structured data instead of raw content:\n\n| Parser | Extracts |\n| --- | --- |\n| `@olostep/amazon-product` | Title, price, rating, reviews, images, variants |\n| `@olostep/google-search` | Result titles, URLs, snippets |\n| `@olostep/google-maps` | Business name, address, rating, reviews |\n| `@olostep/extract-emails` | Email addresses from any page |\n| `@olostep/extract-socials` | Social profile links (X, GitHub, LinkedIn, etc.) |\n| `@olostep/extract-calendars` | Google Calendar and ICS links |\n\nSee the full list in the [Olostep parser store \u2192](https://www.olostep.com/store)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#troubleshooting)  Troubleshooting\n\nAPI key rejected\n\nCopy the key directly from [olostep.com/dashboard](https://olostep.com/dashboard) with no trailing spaces. Delete and recreate the credential in n8n if the error persists.\n\nScraped content is empty\n\nIncrease **Wait Before Scraping** (try 2000\u20135000ms for JS-heavy pages). Confirm the URL is publicly accessible without a login. If a specific domain is consistently failing, contact [info@olostep.com](mailto:info@olostep.com).\n\nBatch URL format error\n\nThe **URLs to Scrape** field expects a JSON array:\n\n```\n[\\\n  { \"url\": \"https://example.com/page-1\", \"custom_id\": \"p1\" },\\\n  { \"url\": \"https://example.com/page-2\", \"custom_id\": \"p2\" }\\\n]\n```\n\nUse a Code node upstream to build this array from your data if needed.\n\nRate limit hit\n\nAdd a **Wait** node between scrape steps, or switch to **Batch Scrape URLs** instead of looping single scrapes. Check current usage in the [dashboard](https://olostep.com/dashboard).\n\nCommunity Nodes not visible in Settings\n\nOn n8n Cloud, community nodes must be enabled by a workspace owner. On self-hosted, make sure `N8N_COMMUNITY_PACKAGES_ENABLED=true` is set in your environment. See [n8n\u2019s installation guide](https://docs.n8n.io/integrations/community-nodes/installation/).\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#related)  Related\n\n[**Scrapes API** \\\\\n\\\\\nFull reference for the scrape endpoint](https://docs.olostep.com/features/scrapes)\n\n[**Batches API** \\\\\n\\\\\nHow batch jobs work and how to retrieve results](https://docs.olostep.com/features/batches)\n\n[**Crawls API** \\\\\n\\\\\nCrawl configuration and result retrieval](https://docs.olostep.com/features/crawls)\n\n[**Maps API** \\\\\n\\\\\nURL discovery and filtering options](https://docs.olostep.com/features/maps)\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#get-started)  Get Started\n\nReady to automate your web search, scraping, and crawling workflows?\n\n[**n8n Website** \\\\\n\\\\\nn8n platform](https://n8n.io/)\n\n[**Install the Node** \\\\\n\\\\\nInstall n8n-nodes-olostep and start building automated workflows](https://www.npmjs.com/package/n8n-nodes-olostep)\n\nConnect Olostep with n8n and automate your web data extraction today!\n\nWas this page helpful?\n\nYesNo",
      "content_chars": 11989,
      "published_date": null
    },
    {
      "rank": 3,
      "url": "https://docs.olostep.com/features/search",
      "title": "Search API - Olostep Docs",
      "content": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/features/search#content-area)\n\nThe Olostep `/v1/searches` endpoint lets you search the web with a natural language query and get back a deduplicated list of relevant links with titles and descriptions.\n\n- Send a query in plain English\n- Get back structured links from across the web\n- Optionally scrape every returned URL in one round-trip and embed `markdown_content` / `html_content` directly into the response\n- Filter by domain, control the result count, and bound the scraping wallclock\n\nIt will search for the query semantically across the web and return results.For API details, see the [Search Endpoint API Reference](https://docs.olostep.com/api-reference/searches/create).\n\n## [\u200b](https://docs.olostep.com/features/search\\#installation)  Installation\n\nPython\n\nNode\n\ncURL\n\nNode (API)\n\nPython (API)\n\n```\npip install olostep\n```\n\n```\nnpm install olostep\n```\n\n```\n# curl is available by default on macOS, Linux, and Windows\n```\n\n```\nnpm install node-fetch\n```\n\n```\npip install requests\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#basic-usage)  Basic usage\n\nSend a natural language query and receive a list of relevant links.\n\nPython\n\nNode\n\ncURL\n\nNode (API)\n\nPython (API)\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.create(\"Best Answer Engine Optimization startups\")\n\nprint(search.id, len(search.links))\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n\nconst search = await client.searches.create('Best Answer Engine Optimization startups')\n\nconsole.log(search.id, search.links.length)\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/searches\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"query\": \"Best Answer Engine Optimization startups\"\n  }'\n```\n\n```\nconst res = await fetch('https://api.olostep.com/v1/searches', {\n  method: 'POST',\n  headers: { 'Authorization': 'Bearer <YOUR_API_KEY>', 'Content-Type': 'application/json' },\n  body: JSON.stringify({\n    query: 'Best Answer Engine Optimization startups'\n  })\n})\nconsole.log(await res.json())\n```\n\n```\nimport requests, json\n\nendpoint = \"https://api.olostep.com/v1/searches\"\npayload = {\n  \"query\": \"Best Answer Engine Optimization startups\"\n}\nheaders = {\"Authorization\": \"Bearer <YOUR_API_KEY>\", \"Content-Type\": \"application/json\"}\n\nresponse = requests.post(endpoint, json=payload, headers=headers)\nprint(json.dumps(response.json(), indent=2))\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#request-parameters)  Request parameters\n\n| Field | Type | Required | Default | Description |\n| --- | --- | --- | --- | --- |\n| `query` | string | yes | \u2014 | The search query in natural language. |\n| `limit` | integer | no | `12` | Maximum number of links to return after deduplication. Must be between `1` and `25`. |\n| `include_domains` | string\\[\\] | no | `[]` | Restrict results to these domains. Bare hosts only \u2014 leading `http(s)://` and trailing slashes are stripped automatically. |\n| `exclude_domains` | string\\[\\] | no | `[]` | Exclude results from these domains. Bare hosts only \u2014 leading `http(s)://` and trailing slashes are stripped automatically. |\n| `scrape_options` | object | no | \u2014 | When provided, every returned link is also scraped and its content embedded in the response. See [scrape\\_options](https://docs.olostep.com/features/search#scrape-options) below. |\n| `fast_mode` | boolean | no | `false` | Request a direct, low-latency search using your query verbatim. Default mode performs a broader search pass for wider result coverage. |\n\n### [\u200b](https://docs.olostep.com/features/search\\#limiting-the-number-of-results)  Limiting the number of results\n\n```\n{\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"limit\": 5\n}\n```\n\n### [\u200b](https://docs.olostep.com/features/search\\#filtering-by-domain)  Filtering by domain\n\n`include_domains` narrows results to a whitelist; `exclude_domains` filters out unwanted sources. They can be combined.\n\n```\n{\n  \"query\": \"OpenAI Sora shutdown analysis\",\n  \"include_domains\": [\"nytimes.com\", \"wsj.com\", \"bbc.com\"],\n  \"exclude_domains\": [\"pinterest.com\"]\n}\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#scrape_options)  scrape\\_options\n\nPass `scrape_options` to scrape every returned URL in parallel and embed the rendered content directly on each link. This saves a round-trip per result vs. calling `/v1/searches` and `/v1/scrapes` separately.\n\n```\n{\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"limit\": 10,\n  \"scrape_options\": {\n    \"formats\": [\"markdown\"],\n    \"remove_css_selectors\": \"default\",\n    \"timeout\": 25\n  }\n}\n```\n\n| Field | Type | Default | Description |\n| --- | --- | --- | --- |\n| `formats` | string\\[\\] | `[\"markdown\"]` | Output formats to attach to each link. For `/v1/searches`, only `\"html\"` and `\"markdown\"` are supported. Pass `[\"html\", \"markdown\"]` to receive both. |\n| `remove_css_selectors` | string | `\"default\"` | Forwarded to `/v1/scrapes`. `\"default\"` strips nav/footer/script/style/svg/dialog noise. Use `\"none\"` to disable, or pass a JSON-stringified array of selectors to remove. |\n| `timeout` | integer | `25` | Wallclock budget in **seconds** for the entire scrape phase. Must be between `1` and `60`. After this elapses, the search returns immediately \u2014 content fields will be `null` for any links that hadn\u2019t finished. |\n\n### [\u200b](https://docs.olostep.com/features/search\\#behavior)  Behavior\n\n- All links are scraped **in parallel**. The `timeout` bounds the whole batch, not each individual link.\n- Per-link scrape failures (network errors, individual page timeouts) leave that link\u2019s `markdown_content` / `html_content` as `null` while other links return normally.\n- If the global `timeout` elapses before all scrapes finish, the search responds immediately with the links it has \u2014 already-completed scrapes keep their content; in-flight ones come back with `null` content.\n- For `reddit.com/.../comments/...` URLs, the request is automatically routed through the `@olostep/reddit-post` parser and the structured JSON is rendered into clean markdown + basic HTML\n- If the combined inline content exceeds 9MB, content fields are nulled, `result.size_exceeded` is set to `true`, and you can fetch the full payload from `result.json_hosted_url`.\n\n### [\u200b](https://docs.olostep.com/features/search\\#example-with-scraping)  Example with scraping\n\nPython\n\nNode\n\ncURL\n\nNode (API)\n\nPython (API)\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.create(\n    query=\"What's going on with OpenAI's Sora shutting down?\",\n    limit=5,\n    scrape_options={\"formats\": [\"markdown\"], \"timeout\": 25},\n)\n\nfor link in search.links:\n    print(link[\"url\"], \"\u2014\", len(link.get(\"markdown_content\") or \"\"), \"chars\")\n```\n\n```\nimport Olostep, { Format } from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n\nconst search = await client.searches.create({\n  query: \"What's going on with OpenAI's Sora shutting down?\",\n  limit: 5,\n  scrapeOptions: {\n    formats: [Format.MARKDOWN],\n    timeout: 25\n  }\n})\n\nfor (const link of search.links) {\n  console.log(link.url, '\u2014', (link.markdown_content || '').length, 'chars')\n}\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/searches\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"query\": \"What'\"'\"'s going on with OpenAI'\"'\"'s Sora shutting down?\",\n    \"limit\": 5,\n    \"scrape_options\": {\n      \"formats\": [\"markdown\"],\n      \"timeout\": 25\n    }\n  }'\n```\n\n```\nconst res = await fetch('https://api.olostep.com/v1/searches', {\n  method: 'POST',\n  headers: { 'Authorization': 'Bearer <YOUR_API_KEY>', 'Content-Type': 'application/json' },\n  body: JSON.stringify({\n    query: \"What's going on with OpenAI's Sora shutting down?\",\n    limit: 5,\n    scrape_options: {\n      formats: ['markdown'],\n      timeout: 25\n    }\n  })\n})\nconst data = await res.json()\nfor (const link of data.result.links) {\n  console.log(link.url, '\u2014', (link.markdown_content || '').length, 'chars')\n}\n```\n\n```\nimport requests, json\n\nendpoint = \"https://api.olostep.com/v1/searches\"\npayload = {\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"limit\": 5,\n  \"scrape_options\": {\n    \"formats\": [\"markdown\"],\n    \"timeout\": 25\n  }\n}\nheaders = {\"Authorization\": \"Bearer <YOUR_API_KEY>\", \"Content-Type\": \"application/json\"}\n\nresponse = requests.post(endpoint, json=payload, headers=headers)\ndata = response.json()\nfor link in data[\"result\"][\"links\"]:\n  print(link[\"url\"], \"\u2014\", len(link.get(\"markdown_content\") or \"\"), \"chars\")\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#response)  Response\n\nYou will receive a `search` object in response. The `search` object contains an `id`, your original `query`, `credits_consumed`, and a `result` with a list of `links`.\n\n```\n{\n  \"id\": \"search_9bi0sbj9xa\",\n  \"object\": \"search\",\n  \"created\": 1760327323,\n  \"metadata\": {},\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"credits_consumed\": 10,\n  \"result\": {\n    \"json_content\": \"...\",\n    \"json_hosted_url\": \"https://olostep-storage.s3.us-east-1.amazonaws.com/search_9bi0sbj9xa.json\",\n    \"size_exceeded\": false,\n    \"credits_consumed\": 10,\n    \"links\": [\\\n      {\\\n        \"url\": \"https://www.bbc.com/news/articles/c3w3e467ewqo\",\\\n        \"title\": \"OpenAI to shut down Sora video platform\",\\\n        \"description\": \"OpenAI says it will discontinue its Sora app...\",\\\n        \"markdown_content\": \"# OpenAI to shut down Sora video platform\\n\\nOpenAI says it will discontinue...\"\\\n      },\\\n      {\\\n        \"url\": \"https://www.reddit.com/r/OutOfTheLoop/comments/1s2u847/whats_going_on_with_openais_sora_shutting_down/\",\\\n        \"title\": \"What's going on with OpenAI's Sora shutting down?\",\\\n        \"description\": \"Reddit thread discussing the shutdown.\",\\\n        \"markdown_content\": \"# What's going on with OpenAI's Sora shutting down?\\n\\n*r/OutOfTheLoop \u00b7 u/rm-minus-r \u00b7 1mo ago*\\n\\n...\"\\\n      }\\\n    ]\n  }\n}\n```\n\nEach link in `result.links` contains:\n\n| Field | Type | Description |\n| --- | --- | --- |\n| `url` | string | The URL of the search result. |\n| `title` | string | The title of the result page. |\n| `description` | string | A short snippet describing the result. |\n| `markdown_content` | string | Markdown content of the page. Only present when `scrape_options.formats` includes `\"markdown\"`. `null` if the scrape failed, was empty, or hit the global timeout. |\n| `html_content` | string | HTML content of the page. Only present when `scrape_options.formats` includes `\"html\"`. `null` on failure/timeout. |\n\nThe full result is also available as a hosted JSON file at `result.json_hosted_url` \u2014 useful when `result.size_exceeded` is `true`.\n\n## [\u200b](https://docs.olostep.com/features/search\\#retrieving-a-past-search)  Retrieving a past search\n\n`GET /v1/searches/{search_id}` returns whatever was persisted at search time, including any scraped content. It\u2019s a pure idempotent read \u2014 no re-scraping, no re-billing. Older searches without `scrape_options` simply have no per-link content fields.\n\nPython\n\nNode\n\ncURL\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.get(search_id=\"search_9bi0sbj9xa\")\nprint(search.id, len(search.links))\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n\nconst search = await client.searches.get('search_9bi0sbj9xa')\nconsole.log(search.id, search.links.length)\n```\n\n```\ncurl -s \"https://api.olostep.com/v1/searches/search_9bi0sbj9xa\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\"\n```\n\nSee [Get Search](https://docs.olostep.com/api-reference/searches/get) for full details.\n\n## [\u200b](https://docs.olostep.com/features/search\\#pricing)  Pricing\n\nEach search costs **5 credits** for the search itself.When `scrape_options` is provided, each scraped page is billed at the standard `/v1/scrapes` rate (typically 1 credit per page; some parsers cost more). The total is returned in `credits_consumed`.Examples:\n\n| Request | `credits_consumed` |\n| --- | --- |\n| Search only | `5` |\n| Search + 5 scraped pages (1 credit each) | `10` |\n\nWas this page helpful?\n\nYesNo",
      "content_chars": 12463,
      "published_date": null
    },
    {
      "rank": 4,
      "url": "https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages",
      "title": "Estimating storage from the number of index pages - IBM",
      "content": "[Skip to content](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages#main-content)[IBM logo](https://www.ibm.com/)Documentation\n[Announcements](https://www.ibm.com/docs/en/announcements)[IBM Redbooks](https://www.ibm.com/docs/en/redbooks)[IBM Product Documentation Directory](https://www.ibm.com/docs/en/products)Table of contentsDark mode\n\nYour current region:\n\n**Unknown \u2013 Unknown**\n\nNo other regions available\n\nMy IBM\nLog in\n\n\nDark mode\n\nDb2 for z/OS\n\nClose table of contents\n\nChange version\n\nSelect\n\n13.0.012.0.011.0.0\n\nShow full table of contents\n\nFilter on titles\n\n- [Welcome to Db2 13 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=welcome-db2-13-zos)\n\n- [About Db2 13 documentation](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=about-db2-13-documentation)\n\n- [Getting started with Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=getting-started-db2-zos)\n\n- [What's new in Db2 13](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=whats-new-in-db2-13)\n\n- [Adopting new capabilities in Db2 13 continuous delivery](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=adopting-new-capabilities-in-db2-13-continuous-delivery)\n\n- [Installing or migrating to Db2 13](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=installing-migrating-db2-13)\n\n- [Db2 application DevOps solutions](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-application-devops-solutions)\n\n- [Administering Db2](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=administering-db2)\n\n\n\n\n\n\n\n  - [Designing and implementing Db2 databases](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-designing-implementing-databases)\n\n\n\n\n\n\n\n    - [Database objects and relationships](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=databases-database-objects-relationships)\n\n    - [Implementing your database design](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=databases-implementing-your-database-design)\n\n\n\n\n\n\n\n      - [Implementing databases](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-databases)\n\n      - [Implementing storage groups](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-storage-groups)\n\n      - [Implementing table spaces](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-table-spaces)\n\n      - [Implementing tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-tables)\n\n      - [Implementing views](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-views)\n\n      - [Implementing indexes](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-indexes)\n\n      - [Implementing schemas](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-schemas)\n\n      - [Loading data into tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-loading-data-into-tables)\n\n      - [implementing stored procedures](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-stored-procedures)\n\n      - [Implementing relationships with referential constraints](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-relationships-referential-constraints)\n\n      - [implementing triggers](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-triggers)\n\n      - [Implementing user-defined functions](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-user-defined-functions)\n\n      - [Implementing Db2 system-defined routines](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-db2-system-defined-routines)\n\n      - [Obfuscating source code of SQL procedures, SQL functions, and triggers](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=iydd-obfuscating-source-code-sql-procedures-sql-functions-triggers)\n\n      - [Estimating disk storage for user data](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-estimating-disk-storage-user-data)\n\n\n\n\n\n\n\n        - [General approach to estimating storage](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=data-general-approach-estimating-storage)\n\n        - [Calculating the space required for a table](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=data-calculating-space-required-table)\n\n        - [Calculating the space required for an index](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=data-calculating-space-required-index)\n\n\n\n\n\n\n\n          - [Levels of index pages](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-levels-pages)\n\n          - [Estimating storage from the number of index pages](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages)\n\n\n      - [Identifying databases that might exceed the OBID limit](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-identifying-databases-that-might-exceed-obid-limit)\n\n\n    - [Altering your database design](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=databases-altering-your-database-design)\n\n\n  - [Operation and recovery](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-operation-recovery)\n\n  - [Exit routines](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-exit-routines)\n\n  - [Data sharing](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-data-sharing)\n\n  - [International data](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-international-data)\n\n  - [Db2 REST services](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-rest-services)\n\n  - [IBM Text Search for Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-text-search-zos)\n\n  - [IBM Spatial Support for Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-spatial-support-zos)\n\n  - [IBM SQL Tuning Services](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-sql-tuning-services)\n\n  - [IBM Db2 for z/OS Agent](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-zos-agent)\n\n\n- [Programming for Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=programming-db2-zos)\n\n- [Securing Db2](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=securing-db2)\n\n- [Managing Db2 performance](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=managing-db2-performance)\n\n- [Running AI queries with SQL Data Insights](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=running-ai-queries-sql-data-insights)\n\n- [Enabling Db2 for IBM Db2 Analytics Accelerator for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=enabling-db2-db2-analytics-accelerator-zos)\n\n- [Implementing Db2 stored procedures](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=implementing-db2-stored-procedures)\n\n- [Db2 SQL](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-sql)\n\n- [Db2 commands](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-commands)\n\n- [Db2 Utilities](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-utilities)\n\n- [Db2 catalog tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-catalog-tables)\n\n- [Db2-supplied user tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-supplied-user-tables)\n\n- [Db2 messages](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-messages)\n\n- [Db2 codes](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-codes)\n\n- [IRLM messages and codes](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=irlm-messages-codes)\n\n- [Troubleshooting problems in Db2](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=troubleshooting-problems-in-db2)\n\n- [Db2 glossary](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-glossary)\n\n- [Notices](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=notices)\n\n[Announcements & sales manuals](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?announcement=all)\n\n[Download PDF](https://www.ibm.com/docs/en/SSEPEK_13.0.0/home/src/tpc/db2z_pdfmanuals.html \"Download PDF\")\n\nOffline docs\n\nFocus sentinel\n\n## Security update required for IBM Docs Offline\n\nClose\n\nA critical security vulnerability has been identified in IBM Docs Offline. You must download and install the latest version to stay protected.\n\nVisit the [Docs Offline page](https://www.ibm.com/docs/en/offline) to download the updated version, or [view the security bulletin](https://www.ibm.com/support/pages/node/7283484 \"Opens in a new tab\") for full details.\n\nGo to Docs Offline pageContinue download\n\nFocus sentinel\n\n[Get hands-on experience with IBM tech\u00a0\u00a0Join one of the largest technical IBM community gatherings!\u00a0\u00a0\u2192](https://www.ibm.com/events/techxchange)\n\nChange version\n\n13.0.012.0.011.0.0\n\nWas this topic helpful?\n\npositive feedback\n\nnegative feedback\n\nFocus sentinel\n\n## Rate this content\n\nClose\n\nGreat! Let us know what you found helpful.\n\nComment 0/750\n\nNote: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support.\n\nCancelSubmit\n\nFocus sentinel\n\nFocus sentinel\n\n## Rate this content\n\nClose\n\nWhat can we do to improve the content?\n\nComment 0/750\n\nNote: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support.\n\nCancelSubmit\n\nFocus sentinel\n\nFocus sentinel\n\n## Provide more feedback\n\nClose\n\nComment 0/750\n\nNote: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support.\n\nCancelSubmit\n\nFocus sentinel\n\nFocus sentinel\n\nClose\n\nRate this content\n\nThank you for your feedback!\n\nTogether, we can continue to improve IBM Documentation.\n\nReturn to topic\n\nFocus sentinel\n\nFocus sentinel\n\nClose\n\n## Thank you for your submission.\n\nSubmissions are limited to 1 per day per topic.\n\nFocus sentinel\n\nFocus sentinel\n\nClose\n\n## Error submitting rating\n\nThere has been an error sending your feedback to the team. Your comment was saved locally, if not in an incognito browser, and will be available when attempting to submit feedback again.\n\nPlease try again later.\n\nFocus sentinel\n\n# Estimating storage from the number of index pages\n\nLast Updated: 2026-01-07\n\nBefore you run a LOAD utility job to load an index, estimate\nthe future storage requirements of the index.\n\n## About this task [Copy to clipboard](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages\\#taskdb2z_estimatestoragefromindex__context__1 \"Copy to clipboard\")\n\nAn index key on an auxiliary table for LOBs is 19 bytes and uses the same formula as other indexes. The RID value that is stored within the index is 4, 5, or 7 bytes, depending on the table space type.\n\nIn general, the\nlength of the index key is the sum of the lengths of all the columns\nof the key, plus the number of columns that allow nulls. The length\nof a varying-length column is the maximum length if the index is padded.\nOtherwise, if an index is not padded, estimate the length of a varying-length\ncolumn to be the average length of the column data, and add a two-byte\nlength field to the estimate. You can retrieve the value of the AVGKEYLEN\ncolumn in the SYSIBM.SYSINDEXES catalog table to determine the average\nlength of keys within an index.\n\nThe following index calculations\nare intended only to help you estimate the storage required for an\nindex. Because there is no way to predict the exact number of duplicate\nkeys that can occur in an index, the results of these calculations\nare not absolute. It is possible, for example, that for a nonunique\nindex, more index entries than the calculations indicate might be\nable to fit on an index page.\n\nImportant: Space allocation\nparameters are specified in kilobytes.\n\nIn the following calculations,\nassume the following:\n\nkThe length of the index key.nThe average number of data records per distinct key value of a\nnonunique index. For example:\n\n- a = number of data records per index\n- b = number of distinct key values per index\n- n = a / b\n\nfThe value of PCTFREE.pThe value of FREEPAGE.rThe record identifier (RID) length. Use 4, 5, or 7 bytes, depending on the table space type:\n\n| Table space type | r value to use |\n| --- | --- |\n| Partition-by-growth (PBG UTS) | 5 |\n| Partition-by-range (PBR UTS) with relative page numbering | 7 |\n| Partition-by-range (PBR UTS) with absolute page numbering | 5 |\n| Non-UTS defined with `DSSIZE 4G` or greater | 5 |\n| Non-UTS defined with `LARGE` | 5 |\n| Other non-UTS types | 4 |\n\nSThe value of the page size minus the length of the page header\nand page tail.FLOORThe operation of discarding the decimal portion of a real number.CEILINGThe operation of rounding a real number up to the next highest\ninteger.MAXThe operation of selecting the highest integer value.\n\n## Procedure [Copy to clipboard](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages\\#taskdb2z_estimatestoragefromindex__steps__1 \"Copy to clipboard\")\n\nTo estimate index storage size, complete the following calculations:\n\n1. Calculate the pages for a unique index.\n1. Calculate the total leaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately k \\+ r \\+ 3\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR((100 - f) \u00d7 S / 100)\n\n      3. Calculate the entries per page\n         entries per page is approximately FLOOR(usable space per page / space per key)\n\n      4. Calculate the **total leaf pages**\n         total leaf pages is approximately CEILING(number of table rows / entries per page)\n\n\n2. Calculate the total nonleaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately k \\+ 7\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR(MAX(90, (100 - f )) \u00d7 S /100)\n\n      3. Calculate the entries per page\n         entries per page is approximately FLOOR(usable space per page / space per key)\n\n      4. Calculate the minimum child pages\n         minimum child pages is approximately MAX(2, (entries per page + 1))\n\n      5. Calculate the level 2 pages\n         level 2 pages is approximately CEILING(total leaf pages / minimum child pages)\n\n      6. Calculate the level 3 pages\n         level 3 pages is approximately CEILING(level 2 pages / minimum child pages)\n\n      7. Calculate the level x pages\n         level x pages is approximately CEILING(previous level pages / minimum child pages)\n\n      8. Calculate the **total nonleaf pages**\n         total nonleaf pages is approximately (level 2 pages \\+ level 3 pages \\+ ... \\+ level x pages until the number of level x pages = 1)\n2. Calculate the pages for a nonunique index.\n1. Calculate the total leaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately 4 + k \\+ (n \u00d7 (r+1))\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR((100 - f ) \u00d7 S / 100)\n\n      3. Calculate the key entries per page\n         key entries per page is approximately n\u00d7 (usable space per page / space per key)\n\n      4. Calculate the remaining space per page\n         remaining space per page is approximately usable space per page \\- (key entries per page / n) \u00d7space per key\n\n      5. Calculate the data records per partial entry\n         data records per partial entry is approximately FLOOR((remaining space per page \\- (k \\+ 4)) / 5)\n\n      6. Calculate the partial entries per page\n         partial entries per page is approximately (n / CEILING(n / data records per partial entry)) if data records per partial entry >= 1, or 0 if data records per partial entry < 1\n\n      7. Calculate the entries per page\n         entries per page is approximately MAX(1, (key entries per page \\+ partial entries per page))\n\n      8. Calculate the **total leaf pages**\n         total leaf pages is approximately CEILING(number of table rows / entries per page)\n\n\n2. Calculate the total nonleaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately k \\+ r \\+ 7\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR (MAX(90, (100- f))\u00d7 S / 100)\n\n      3. Calculate the entries per page\n         entries per page is approximately FLOOR((usable space per page / space per key)\n\n      4. Calculate the minimum child pages\n         minimum child pages is approximately MAX(2, (entries per page \\+ 1))\n\n      5. Calculate the level 2 pages\n         level 2 pages is approximately CEILING(total leaf pages / minimum child pages)\n\n      6. Calculate the level 3 pages\n         level 3 pages is approximately CEILING(level 2 pages / minimum child pages)\n\n      7. Calculate the level x pages\n         level x pages is approximately CEILING(previous level pages / minimum child pages)\n\n      8. Calculate the **total nonleaf pages**\n         total nonleaf pages is approximately (level 2 pages \\+ level 3 pages \\+ ... \\+ level x pages until x = 1)\n3. Calculate the pages for an index that is not compressed.\n1. Calculate the usable space per leaf page:\n\n\n      usable space per leaf page is approximately FLOOR((100 - f) \u00d7 S / 100)\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 62 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n2. Calculate the usable space per nonleaf page:\n\n\n      usable space per nonleaf page is approximately FLOOR (MAX (90, (100 - f ) ) \u00d7 S / 100)\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 48 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n3. Calculate the usable space per space map:\n\n\n      usable space per space map is approximately CEILING ( (tree pages \\+ free pages) / S), where S equals (page size \u2212 header length \u2212 tail length) \u00d7 2 \u2212 1.\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 28 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n4. Calculate the pages for a compressed index.\n1. Calculate the usable space per leaf page:\n\n\n      usable space per leaf page is approximately FLOOR((100 - f) \u00d7 S / 100)\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 66 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n2. Calculate the usable space per nonleaf page:\n\n\n      usable space per nonleaf page is approximately FLOOR (MAX (90, (100 - f ) ) \u00d7 S / 100)\n\n\n\n      The page size is 4096 bytes for 4 KB, 8 KB, 16 KB, and 32 KB page sizes. The length of the page header is 48 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n3. Calculate the usable space per space map:\n\n\n      usable space per space map is approximately CEILING ( (tree pages + free pages) / S), where S equals (page size \u2212 header length \u2212 tail length) \u00d7 2 \u2212 1.\n\n\n\n      The page size is 4096 bytes for 4 KB, 8 KB, 16 KB, and 32 KB page sizes. The length of the page header is 28 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n5. Calculate the total space requirement by estimating the number of kilobytes required for an index built by the LOAD utility.\n1. Calculate the free pages\n\n\n      free pages is approximately FLOOR(total leaf pages / p), or 0 if p = 0\n\n2. Calculate the space map pages\n\n\n      space map pages is approximately CEILING((tree pages \\+ free pages) / S)\n\n3. Calculate the tree pages\n\n\n      tree pages is approximately MAX(2, (total leaf pages \\+ total nonleaf pages))\n\n4. Calculate the total index pages\n\n\n      total index pages is approximately MAX(4, (1 + tree pages \\+ free pages \\+ space map pages))\n\n5. Calculate the **total space requirement**\n\n\n      total space requirement is approximately 4 \u00d7 (total index pages \\+ 2)\n\n## Example [Copy to clipboard](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages\\#taskdb2z_estimatestoragefromindex__example__1 \"Copy to clipboard\")\n\nIn the following example of the entire calculation, assume that an index is defined with these characteristics:\n\n- The index is unique.\n- The table it indexes has 100000 rows.\n- The key is a single column defined as CHAR(10) NOT NULL.\n- The value of PCTFREE is 5.\n- The value of FREEPAGE is 4.\n- The page size is 4 KB.\n\n| Quantity | Calculation | Result |\n| --- | --- | --- |\n| Length of key<br>Average number of duplicate keys<br>PCTFREE<br>FREEPAGE | k<br>n<br>f<br>p | 10<br>1<br>5<br>4 |\n| **Calculate total leaf pages**<br>Space per key<br>Usable space per page<br>Entries per page<br>Total leaf pages | k \\+ 7<br>FLOOR((100 - f ) \u00d7 4032/100)<br>FLOOR(usable space per page / space per key)<br>CEILING(number of table rows / entries per page) | 17<br>3844<br>225<br>445 |\n| **Calculate total nonleaf pages**<br>Space per key<br>Usable space per page<br>Entries per page<br>Minimum child pages<br>Level 2 pages<br>Level 3 pages<br>Total nonleaf pages | k \\+ 7<br>FLOOR(MAX(90, (100 - f )) \u00d7 4046/100)<br>FLOOR(usable space per page / space per key)<br>MAX(2, (entries per page \\+ 1))<br>CEILING(total leaf pages / minimum child pages)<br>CEILING(level 2 pages / minimum child pages)<br>(level 2 pages \\+ level 3 pages \\+\u2026\\+ level x pages until x = 1) | 17<br>3843<br>226<br>227<br>2<br>1<br>3 |\n| **Calculate total space required**<br>Free pages<br>Tree pages<br>Space map pages<br>Total index pages<br>TOTAL SPACE REQUIRED, in KB | FLOOR(total leaf pages / p), or 0 if p = 0<br>MAX(2, (total leaf pages \\+ total nonleaf pages))<br>CEILING((tree pages \\+ free pages)/8131)<br>MAX(4, (1 + tree pages \\+ free pages \\+ space map pages))<br>4 \u00d7 (total index pages \\+ 2) | 111<br>448<br>1<br>561<br>2252 |\n\nTable 1. Sample of the total space requirement for a unique index\n\nWhile IBM values the use of inclusive language, terms that are outside of IBM's direct influence, for the sake of maintaining user understanding, are sometimes required. As other industry leaders join IBM in embracing the use of inclusive language, IBM will continue to update the documentation to reflect those changes.\n\n\u00a9 Copyright IBM Corporation 1983, 2026\n\n[IBM logo](https://www.ibm.com/)\n\nArabic / \u0639\u0631\u0628\u064a\u0629\n\nBulgarian / \u0411\u044a\u043b\u0433\u0430\u0440\u0441\u043a\u0438\n\nCatalan / Catal\u00e0\n\nCzech / \u010ce\u0161tina\n\nDanish / Dansk\n\nGerman / Deutsch\n\nGreek / \u0395\u03bb\u03bb\u03b7\u03bd\u03b9\u03ba\u03ac\n\nEnglish\n\nSpanish / Espa\u00f1ol\n\nFinnish / Suomi\n\nFrench / Fran\u00e7ais\n\nCroatian / Hrvatski\n\nHungarian / Magyar\n\nItalian / Italien\n\nHebrew / \u05e2\u05d1\u05e8\u05d9\u05ea\n\nJapanese / \u65e5\u672c\u8a9e\n\nKorean / \ud55c\uad6d\uc5b4\n\nKazakh / \u049a\u0430\u0437\u0430\u049b\u0448\u0430\n\nDutch / Nederlands\n\nNorwegian / Norsk\n\nPolish / polski\n\nPortuguese/Brazil / Portugu\u00eas/Brasil\n\nPortuguese/Portugal / Portugu\u00eas/Portugal\n\nRomanian / Rom\u00e2n\u0103\n\nRussian / \u0420\u0443\u0441\u0441\u043a\u0438\u0439\n\nSlovak / Sloven\u010dina\n\nSlovenian / sloven\u0161\u010dina\n\nSerbian / srpski\n\nSwedish / Svenska\n\nThai / \u0e20\u0e32\u0e29\u0e32\u0e44\u0e17\u0e22\n\nTurkish / T\u00fcrk\u00e7e\n\nVietnamese / Vi\u00ea\u0323t\n\nChinese Simplified / \u7b80\u4f53\u4e2d\u6587\n\nChinese Traditional / \u7e41\u9ad4\u4e2d\u6587[IBM logo](https://www.ibm.com/)Contact IBMPrivacyTerms of useAccessibilityCookie Preferences\n\nArabic / \u0639\u0631\u0628\u064a\u0629\n\nBulgarian / \u0411\u044a\u043b\u0433\u0430\u0440\u0441\u043a\u0438\n\nCatalan / Catal\u00e0\n\nCzech / \u010ce\u0161tina\n\nDanish / Dansk\n\nGerman / Deutsch\n\nGreek / \u0395\u03bb\u03bb\u03b7\u03bd\u03b9\u03ba\u03ac\n\nEnglish\n\nSpanish / Espa\u00f1ol\n\nFinnish / Suomi\n\nFrench / Fran\u00e7ais\n\nCroatian / Hrvatski\n\nHungarian / Magyar\n\nItalian / Italien\n\nHebrew / \u05e2\u05d1\u05e8\u05d9\u05ea\n\nJapanese / \u65e5\u672c\u8a9e\n\nKorean / \ud55c\uad6d\uc5b4\n\nKazakh / \u049a\u0430\u0437\u0430\u049b\u0448\u0430\n\nDutch / Nederlands\n\nNorwegian / Norsk\n\nPolish / polski\n\nPortuguese/Brazil / Portugu\u00eas/Brasil\n\nPortuguese/Portugal / Portugu\u00eas/Portugal\n\nRomanian / Rom\u00e2n\u0103\n\nRussian / \u0420\u0443\u0441\u0441\u043a\u0438\u0439\n\nSlovak / Sloven\u010dina\n\nSlovenian / sloven\u0161\u010dina\n\nSerbian / srpski\n\nSwedish / Svenska\n\nThai / \u0e20\u0e32\u0e29\u0e32\u0e44\u0e17\u0e22\n\nTurkish / T\u00fcrk\u00e7e\n\nVietnamese / Vi\u00ea\u0323t\n\nChinese Simplified / \u7b80\u4f53\u4e2d\u6587\n\nChinese Traditional / \u7e41\u9ad4\u4e2d\u6587\n\n[![close icon](https://consent.trustarc.com/get?name=ibm_close_icon.svg)](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages#)\n\nIBM web domains\n\nibm.com, ibm.org, ibm-zcouncil.com, insights-on-business.com, jazz.net, mobilebusinessinsights.com, promontory.com, proveit.com, ptech.org, s81c.com, securityintelligence.com, skillsbuild.org, softlayer.com, storagecommunity.org, think-exchange.com, thoughtsoncloud.com, alphaevents.webcasts.com, ibm-cloud.github.io, ibmbigdatahub.com, bluemix.net, mybluemix.net, ibm.net, ibmcloud.com, galasa.dev, blueworkslive.com, swiss-quantum.ch, blueworkslive.com, cloudant.com, ibm.ie, ibm.fr, ibm.com.br, ibm.co, ibm.ca, community.watsonanalytics.com, datapower.com, skills.yourlearning.ibm.com, bluewolf.com, carbondesignsystem.com, openliberty.io\n\n![close icon](https://consent.trustarc.com/get?name=ibm_close_icon.svg)\n\nAbout cookies on this siteOur websites require some cookies to function properly (required). In addition, other cookies may be used with your consent to analyze site usage, improve the user experience and for advertising.For more information, please review your cookie\u00a0preferences\u00a0options. By visiting our website, you agree to our processing of information as described in IBM\u2019s [privacy\u00a0statement](https://www.ibm.com/privacy).\u00a0 To provide a smooth navigation, your cookie preferences will be shared across the IBM web domains listed\u00a0[here](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages#truste_domain_list).\n\nAccept AllMore options\n\n![](https://analytics.o11.tech/e/a.gif?aqet=pv&evid=e3ba4354-61ac-487e-a753-01d0414b46d7&aq_m=1&pubid=b0a6ae0081d4f8392fc41064e9ae7fb9&dmn=www.ibm.com&tt=tcs.dhj&cid=c076&lbl=c076&flbl=pxcel&ll=e&ver=1.2115.265&ell=e&cck=_autid&pn=%2Fdocs%2Fen%2Fdb2-for-zos%2F13.0.0&qs=topic%3Dindex-estimating-storage-from-number-pages&rdn=www.google.com&rpn=%2F&rqs=na&cc=US&cont=NA&rc=VA&urls=!1!0!b-13w,!1!0!b-13x,!1!0!b-13y,!1!0!b-144,!1!0!b-14i&rnd=1788514043583&cid=c076&version=1.2115.265&cc=US&cont=NA&repeat=0&htmLcy=136)\n\n![](https://a.usbrowserspeed.com/cs?pid=8046b079bd55b4d00035e43c6d9f73c385a7184438a37f8e4401bf0b831c02f6&puid=4908fe26891c9b4c6351c2cee71801d1;e3ba4354-61ac-487e-a753-01d0414b46d7;6a9a8efb8e6bdc174b1d04b8&r=https%3A%2F%2Fanalytics.o11.tech%2F1%2Fe%2Fc.gif%3Faqet%3Didsync%26puu%3D%24%7BDEVICE_ID%7D)\n\n![](https://analytics.o11.tech/e/a.gif?aqet=pv&evid=e3ba4354-61ac-487e-a753-01d0414b46d7&aq_m=1&pubid=4908fe26891c9b4c6351c2cee71801d1&dmn=www.ibm.com&tt=tcs.dhj&cid=c076&lbl=c076&flbl=pxcel&ll=e&ver=1.2115.265&ell=e&cck=_autid&pn=%2Fdocs%2Fen%2Fdb2-for-zos%2F13.0.0&qs=topic%3Dindex-estimating-storage-from-number-pages&rdn=www.google.com&rpn=%2F&rqs=na&cc=US&cont=NA&rc=VA&urls=!1!0!b-13w,!1!0!b-13x,!1!0!b-13y,!1!0!b-144,!1!0!b-14i&rnd=1788514043563&cid=c076&version=1.2115.265&cc=US&cont=NA&repeat=0&htmLcy=110)\n\n![](https://a.usbrowserspeed.com/cs?pid=8046b079bd55b4d00035e43c6d9f73c385a7184438a37f8e4401bf0b831c02f6&puid=32c5475300e9fc6db3d3688149e141ea;e3ba4354-61ac-487e-a753-01d0414b46d7;6a9a8efb8e6bdc174b1d04b8&r=https%3A%2F%2Fanalytics.o11.tech%2F1%2Fe%2Fc.gif%3Faqet%3Didsync%26puu%3D%24%7BDEVICE_ID%7D)\n\n![](https://analytics.o11.tech/e/a.gif?aqet=pv&evid=e3ba4354-61ac-487e-a753-01d0414b46d7&aq_m=1&pubid=32c5475300e9fc6db3d3688149e141ea&dmn=www.ibm.com&tt=tcs.dhj&cid=c076&lbl=c076&flbl=pxcel&ll=e&ver=1.2115.265&ell=e&cck=_autid&pn=%2Fdocs%2Fen%2Fdb2-for-zos%2F13.0.0&qs=topic%3Dindex-estimating-storage-from-number-pages&rdn=www.google.com&rpn=%2F&rqs=na&cc=US&cont=NA&rc=VA&urls=!1!0!b-13w,!1!0!b-13x,!1!0!b-13y,!1!0!b-144,!1!0!b-14i&rnd=1788514043600&cid=c076&version=1.2115.265&cc=US&cont=NA&repeat=0&htmLcy=159)",
      "content_chars": 27865,
      "published_date": null
    },
    {
      "rank": 5,
      "url": "https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918",
      "title": "Fetching all docs in an app search index - Elastic Discuss",
      "content": "[Skip to where you left off (last reply, post 6)](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/6) [Skip to top](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/1)\n\n[Skip to main content](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918#main-container)\n\n# [Fetching all docs in an app search index](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918)\n\n[Elastic Search](https://discuss.elastic.co/c/search/84)\n\n- [elastic-app-search](https://discuss.elastic.co/tag/elastic-app-search/118)\n\nYou have selected **0** posts.\n\n[select all](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918)\n\n[cancel selecting](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918)\n\n2.1k\nviews\n\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)3](https://discuss.elastic.co/u/frankjoh2 \"frankjoh2\")\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)2](https://discuss.elastic.co/u/goodroot \"goodroot\")\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/1 \"Jump to the first post\")\n\n6 / 6\n\n\nJan 2019\n\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/6)\n\n## post by frankjoh2 on Dec 7, 2018\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)](https://discuss.elastic.co/u/frankjoh2)\n\n[frankjoh2](https://discuss.elastic.co/u/frankjoh2)\n\n1\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918 \"Post date\")\n\nHi,\n\nI'm trying to get the ids of all documents in my app search index. I tried to simply iteration of searches with an empty query and incrementally increasing the page-number. This seemed to work fine for the first 10 requests.\n\nThis is what my requests looks like:\n\n{\n\n\"query\": \"\",\n\n\"result\\_fields\": {\n\n\"id\": { \"raw\": {} }\n\n},\n\n\"page\": {\"size\": 1000,\"current\": **\\[PAGENUMER\\]** }\n\n}\n\nWhere **\\[PAGENUMBER\\]** is 1 for the first request, 2 for second and so on\u2026\n\nThis is the result of the 10th request:\n\n_{_\n\n_\"meta\": {_\n\n_\"warnings\": \\[\\],_\n\n_\"page\": {_\n\n_\"current\": 10,_\n\n_\"total\\_pages\": 92,_\n\n_\"total\\_results\": 91035,_\n\n_\"size\": 1000_\n\n_},_\n\n_\"request\\_id\": \"39684c716fe14725a70406a1a71789e4\"_\n\n_},_\n\n_\"results\": \\[...\\] <--- 1000 results here_\n\n_}_\n\nWorking as expected and showing all 91035 docs in the index and that there are a total of 92 pages.\n\nBut the result of the 11th request:\n\n_{_\n\n_\"meta\": {_\n\n_\"warnings\": \\[\\],_\n\n_\"page\": {_\n\n_\"current\": 11,_\n\n_\"total\\_pages\": 0,_\n\n_\"total\\_results\": 0,_\n\n_\"size\": 1000_\n\n_},_\n\n_\"request\\_id\": \"39684c716fe14725a70406a1a71789e4\"_\n\n_},_\n\n_\"results\": \\[\\] <--- 0 results here_\n\n_}_\n\nSuddenly it indicates that there are no docs found at all\u2026\n\nThe documentation says that search-request should support up to 1000 in page size and up to 500 pages, but it seems like it supports max 10 pages when page-size is 1000. Or is there some setting I need to change to support more result-pages?\n\nOr is there some other way I can request ids of all docs in the index?\n\nAny help would be appreciated\n\n2.1k\nviews\n\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)3](https://discuss.elastic.co/u/frankjoh2 \"frankjoh2\")\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)2](https://discuss.elastic.co/u/goodroot \"goodroot\")\n\n## post by frankjoh2 on Dec 7, 2018\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)](https://discuss.elastic.co/u/frankjoh2)\n\n[frankjoh2](https://discuss.elastic.co/u/frankjoh2)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/2 \"Post date\")\n\nI found out now that there is a list-method (/documents/list) this is specifically meant to get all docs in an index. But this will return all data (not just the id-field) and has a max pagesize of 100 docs.\n\nThis means I'll have to spend at least 10 times as many api-requests to get this done and that might push me above the monthly limit and resulting in extra licensing-costs.\n\nSo I'd still be interested to know of any alternatives if they exist.\n\n## post by goodroot on Dec 7, 2018\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)](https://discuss.elastic.co/u/goodroot)\n\n[goodroot](https://discuss.elastic.co/u/goodroot)[Kellen Evan](https://discuss.elastic.co/u/goodroot)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/3 \"Post date\")\n\nFrank!\n\nYou are correct. The most effective method of returning documents would be to iterate over the `documents/list` endpoint, like so:\n\n```rust\n\ncurl -X GET 'https://host-xxxxxx.api.swiftype.com/api/as/v1/engines/example-engine/documents/list' \\\n-H 'Content-Type: application/json' \\\n-H 'Authorization: Bearer private-xxxxxxxxxxxxxxxxxxxx' \\\n-d '{\n  \"page\": {\n    \"current\": 1,\n    \"size\": 100\n  }\n}'\n```\n\nAs you pointed out, this will return full documents, not just the `id`. You will need to parse out the `id` when assembling your list.\n\nThanks for posting, I wish you an excellent end to your week.\n\n## post by frankjoh2 on Dec 10, 2018\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)](https://discuss.elastic.co/u/frankjoh2)\n\n[frankjoh2](https://discuss.elastic.co/u/frankjoh2)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/4 \"Post date\")\n\nHi goodroot and thanks for the reply. I tried to use the documents/list endpoint now and sadly that doesn't work either. When the \"current\"-attribute in the request gets higher than 100 the response always returns the 100th page. In other words, it's not possible to get more than the first 10.000 documents (100 pages with 100 documents each) with this endpoint too.\n\n## post by goodroot on Dec 10, 2018\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)](https://discuss.elastic.co/u/goodroot)\n\n[goodroot](https://discuss.elastic.co/u/goodroot)[Kellen Evan](https://discuss.elastic.co/u/goodroot)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/5 \"Post date\")\n\nFrank --\n\nI poked around to try and find you a better answer, but that is correct.\n\nThe limit of both `current` and `size` is `100`, and the endpoint cannot be used to return more than `10,000` documents.\n\nI understand this creates a gap. The documents, and their ids, are available by querying or through the documents dashboard view, but that isn't helpful to one looking to generate a comprehensive list. ![:confused:](https://emoji.discourse-cdn.com/twitter/confused.png?v=6)\n\nThe limit may rise in the future, but as of now it is kept restrictive. If this is a major blocker in your use-case, please email [support@swiftype.com](mailto:support@swiftype.com), referencing this ticket so that we can learn more.\n\nEnjoy the week,\n\nKellen\n\n28 days later\n\n\n## Closed on Jan 7, 2019\n\n[![](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png)](https://discuss.elastic.co/u/system)\n\nClosed on Jan 7, 2019\n\nThis topic was automatically closed 28 days after the last reply. New replies are no longer allowed.\n\nReply\n\n### Related topics\n\n| Topic | Replies | Views | Activity |\n| --- | --- | --- | --- |\n| [Getting next 10k documents with AppSearch.list\\_documents()](https://discuss.elastic.co/t/getting-next-10k-documents-with-appsearch-list-documents/345216)<br>[Elasticsearch](https://discuss.elastic.co/c/elastic-stack/elasticsearch/6) | [5](https://discuss.elastic.co/t/getting-next-10k-documents-with-appsearch-list-documents/345216/1) | 254 | [Oct 2023](https://discuss.elastic.co/t/getting-next-10k-documents-with-appsearch-list-documents/345216/6) |\n| [Results per Page limit](https://discuss.elastic.co/t/results-per-page-limit/217643)<br>[Elastic Search](https://discuss.elastic.co/c/search/84) <br>- [elastic-app-search](https://discuss.elastic.co/tag/elastic-app-search/118) | [5](https://discuss.elastic.co/t/results-per-page-limit/217643/1) | 573 | [Feb 2020](https://discuss.elastic.co/t/results-per-page-limit/217643/6) |\n| [Get all ids with Python](https://discuss.elastic.co/t/get-all-ids-with-python/344689)<br>[Elasticsearch](https://discuss.elastic.co/c/elastic-stack/elasticsearch/6) | [1](https://discuss.elastic.co/t/get-all-ids-with-python/344689/1) | 394 | [Oct 2023](https://discuss.elastic.co/t/get-all-ids-with-python/344689/2) |\n| [How to get results over 10K in App search](https://discuss.elastic.co/t/how-to-get-results-over-10k-in-app-search/351310)<br>[Elastic Search](https://discuss.elastic.co/c/search/84) <br>- [elastic-app-search](https://discuss.elastic.co/tag/elastic-app-search/118) | [10](https://discuss.elastic.co/t/how-to-get-results-over-10k-in-app-search/351310/1) | 13.3k | [Jan 2024](https://discuss.elastic.co/t/how-to-get-results-over-10k-in-app-search/351310/11) |\n| [I need to fetch all document but this query only return 10 documents. I\u2019m checking through postman](https://discuss.elastic.co/t/i-need-to-fetch-all-document-but-this-query-only-return-10-documents-im-checking-through-postman/229574)<br>[Elasticsearch](https://discuss.elastic.co/c/elastic-stack/elasticsearch/6) | [1](https://discuss.elastic.co/t/i-need-to-fetch-all-document-but-this-query-only-return-10-documents-im-checking-through-postman/229574/1) | 1.1k | [Apr 2020](https://discuss.elastic.co/t/i-need-to-fetch-all-document-but-this-query-only-return-10-documents-im-checking-through-postman/229574/2) |\n\nTopic list, column headers with buttons are sortable.\n\n\u00a9 2020\\. All Rights Reserved - Elasticsearch\n\n- Elasticsearch is a trademark of Elasticsearch BV, registered in the U.S.\nand in other countries\n\n- [Trademarks](https://www.elastic.co/legal/trademarks)\n- [Terms](https://www.elastic.co/legal/terms-of-use)\n- [Privacy](https://www.elastic.co/legal/privacy-policy)\n- [Brand](https://www.elastic.co/brand)\n- [Code of Conduct](https://www.elastic.co/community/codeofconduct)\n\nApache, Apache Lucene, Apache Hadoop, Hadoop, HDFS and the yellow elephant\nlogo are trademarks of the\n[Apache Software Foundation](http://www.apache.org/)\nin the United States and/or other\u00a0countries.",
      "content_chars": 10333,
      "published_date": null
    },
    {
      "rank": 6,
      "url": "https://medium.com/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603",
      "title": "Turning the Web Into a Real-Time Database with OloStep - Medium",
      "content": "[Sitemap](https://medium.com/sitemap/sitemap.xml)\n\n[Open in app](https://play.google.com/store/apps/details?id=com.medium.reader&referrer=utm_source%3DmobileNavBar&source=---top_nav_layout_nav-----------------------------------------)\n\nSign up\n\n[Sign in](https://medium.com/m/signin?operation=login&redirect=https%3A%2F%2Fmedium.com%2F%40davidfagb%2Fturning-the-web-into-a-real-time-database-with-olostep-f175e956a603&source=post_page---top_nav_layout_nav-----------------------global_nav------------------)\n\n[Medium Logo](https://medium.com/?source=---top_nav_layout_nav-----------------------------------------)\n\nGet app\n\n[Write](https://medium.com/m/signin?operation=register&redirect=https%3A%2F%2Fmedium.com%2Fnew-story&source=---top_nav_layout_nav-----------------------new_post_topnav------------------)\n\n[Search](https://medium.com/search?source=---top_nav_layout_nav-----------------------------------------)\n\nSign up\n\n[Sign in](https://medium.com/m/signin?operation=login&redirect=https%3A%2F%2Fmedium.com%2F%40davidfagb%2Fturning-the-web-into-a-real-time-database-with-olostep-f175e956a603&source=post_page---top_nav_layout_nav-----------------------global_nav------------------)\n\n![Unknown user](https://miro.medium.com/v2/resize:fill:32:32/1*dmbNkD5D-u45r44go_cf0g.png)\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:40:40/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---post_author_sidebar--f175e956a603-----------------d1058d8a6759----------------------)\n\n## David Fagbuyiro\n\nTechnical writer\n\nFollow writer\n\n[Web Scraping](https://medium.com/tag/web-scraping?source=post_page---header_tags--f175e956a603---------------------------------------)\n\n[Database](https://medium.com/tag/database?source=post_page---header_tags--f175e956a603---------------------------------------)\n\n[Web](https://medium.com/tag/web?source=post_page---header_tags--f175e956a603---------------------------------------)\n\n# Turning the Web Into a Real-Time Database with OloStep\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:32:32/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---byline--f175e956a603---------------------------------------)\n\n[David Fagbuyiro](https://medium.com/@davidfagb?source=post_page---byline--f175e956a603---------------------------------------)\n\nFollow\n\n4 min read\n\n\u00b7\n\nOct 8, 2025\n\n3\n\n[Listen](https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2Fplans%3Fdimension%3Dpost_audio_button%26postId%3Df175e956a603&operation=register&redirect=https%3A%2F%2Fmedium.com%2F%40davidfagb%2Fturning-the-web-into-a-real-time-database-with-olostep-f175e956a603&source=---header_actions--f175e956a603---------------------post_audio_button------------------)\n\nShare\n\nYou\u2019re building an AI assistant for a startup CEO who needs to monitor competitors, track pricing updates, new hires, and product launches. Straightforward stuff. So, you start hunting for data scraping solutions that are too brittle to break after a minor site change. Most sites do not offer APIs. Google? Buried in ads, captchas, and unstructured results. You hit a wall.\n\nThen it hits you that the data is all out there, it\u2019s just trapped inside messy webpages. That\u2019s precisely what OloStep is built to fix. Instead of scraping, crawling, or building static indexes, OloStep flips the model. It treats the open web itself as a live, structured, and queryable database in real time.\n\nPress enter or click to view image in full size\n\n![Illustration of a chaotic web full of HTML pages turning into a clean, structured database view](https://miro.medium.com/v2/resize:fit:700/1*kCHsAMclE1H0wYXNZjqlIA.png)\n\n### The Web Wasn\u2019t Meant to Be Queried\n\nThe internet is one of the most extensive and valuable datasets in history. However, it was not designed for structure, but it was created for publishing.\n\nThere are billions of pages, including blogs, product listings, pricing tables, and job boards, which are constantly being updated. Yet, querying this information in a traditional manner is nearly impossible. You can crawl, index, and scrape the web, but these methods are often brittle, slow, and quickly become outdated.\n\nWhat if, instead, you could treat the open web as if it were your live backend?\n\n## What is Olostep?\n\nOlostep is a powerful web scraping API that efficiently provides e data from any website. It meets the essential demands for fast, reliable, and cost-effective data acquisition, AI development and large-scale data aggregation. rtups, and established companies enhance their data-driven applications and automate workflows.\n\n### The Core Idea: Query the Open Web Like a Database\n\nThink SQL, but the tables are web pages. Think APIs, but the endpoints are search queries. OloStep turns the open web into a real-time knowledge source structured, filterable, and directly useful for agents, workflows, and apps.\n\nLet\u2019s say you\u2019re building an AI agent that finds pricing trends for B2B SaaS tools. Traditionally, you\u2019d:\n\n- Crawl pages\n- Parse them manually\n- Store them\n- Run queries later\n\nBut with OloStep:\n\n```\nSELECT price, plan_name, vendor\nFROM web\nWHERE category = \"CRM software\" AND updated_within = \"7 days\"\n```\n\nIt\u2019s not a stretch. This is the direction we\u2019re heading to obtain structured results directly from the public internet, without scraping or even relying on pipelines, just answers.\n\nPress enter or click to view image in full size\n\n![Side-by-side comparison\u200a\u2014\u200aLeft: messy HTML page, Right: clean JSON-like table extracted live].](https://miro.medium.com/v2/resize:fit:700/1*NL3okKxPHQ4mnYe52GbmvQ.png)\n\nSide-by-side comparison \u2014 Left: messy HTML page, Right: clean JSON-like table extracted live\\].\n\n### What Makes OloStep Different\n\nMost tools that extract data from the web fall into one of two buckets:\n\n1. **Scrapers**: Fast but fragile. Break often, need constant maintenance.\n2. **Knowledge graphs:** Structured but stale. Require time-consuming ingestion pipelines.\n\nOloStep is different because it utilizes AI-native tools to extract structured data pages in real **-** time. The model understands both layout and meaning. It can:\n\n- Recognize product specs on landing pages\n- Pull out job listings across multiple sites\n- Extract facts from articles or documentation\n\nIt\u2019s a web-native query engine. No scraping rules. No brittle XPath selectors. If you\u2019re building anything that depends on real-world knowledge, you don\u2019t want to be stuck waiting for someone else to scrape, ingest, and update it for you.\n\nPress enter or click to view image in full size\n\n![Flowchart showing an AI agent querying OloStep -> real-time web pages -> structured results](https://miro.medium.com/v2/resize:fit:700/1*jl9B2mmsvoodZF5HbGzcoA.png)\n\nFlowchart showing an AI agent querying OloStep -> real-time web pages -> structured results\n\n### Why Now? Because the Stack Just Clicked\n\nThis wasn\u2019t possible five years ago.\n\n- **LLMs** were too weak.\n- **Scrapers** were too brittle.\n- **Knowledge graphs** were too slow.\n\nHowever, now LLMs can comprehend, reason over tables, and infer meaning across web pages in real-time.\n\n## Get David Fagbuyiro\u2019s stories in\u00a0your\u00a0inbox\n\nJoin Medium for free to get updates from\u00a0this\u00a0writer.\n\nSubscribe\n\nSubscribe\n\nRemember me for faster sign in\n\nOloStep rides that wave, turning raw web content into structured, semantically rich data at the moment you need it. It\u2019s not just search. It\u2019s understanding **.**\n\n## Use Cases Already Emerging\n\nBelow are the benefits of implementing OloStep:\n\n- AI Research Agents: agents that answer questions by citing live data from across the web, not just summaries from cached pages.\n- Competitive Intelligence: pull pricing changes, product launches, or staffing shifts from public pages and career portals.\n- Custom Tools & Dashboards: develop internal and public-facing information, such as public tenders, vendor updates, or compliance announcements.\n- Programmatic Search: replace brittle Google scraping with structured queries and filters. No hacks. No captchas. Just answers.\n\n### The Future: Don\u2019t Store the Web. Ask It.\n\nOloStep flips the traditional methods so instead of trying to tame the chaos of the internet into a static database, it embraces the mess and makes it readable, understandable, and usable.\n\nThink of it like this:\n\n- Search gives you links.\n- Scraping gives you fragments.\n- OloStep provides you with facts, and it does so in real-time.\n\nSo instead of building brittle scraping tools or waiting for someone to publish a dataset, you can now ask the web a question and get structured, filtered, useful data in return.\n\nThe web is no longer just something you read; it\u2019s something you query.\n\n### Try It Yourself\n\nIf you\u2019re building agents, dashboards, or any product that depends on fresh, structured information from the real world, OloStep gives you the edge.\n\nCheck it out at [https://olostep.com](https://olostep.com/)\n\nStart asking the web real questions and get real answers.\n\nWith OloStep, you don\u2019t have to.\n\n[Web Scraping](https://medium.com/tag/web-scraping?source=post_page---footer_tags--f175e956a603---------------------------------------)\n\n[Database](https://medium.com/tag/database?source=post_page---footer_tags--f175e956a603---------------------------------------)\n\n[Web](https://medium.com/tag/web?source=post_page---footer_tags--f175e956a603---------------------------------------)\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:48:48/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:64:64/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\nFollow\n\n[**Written by David Fagbuyiro**](https://medium.com/@davidfagb?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\n[54 followers](https://medium.com/@davidfagb/followers?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\n\u00b7 [6 following](https://medium.com/@davidfagb/following?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\nTechnical writer\n\nFollow\n\n[Help](https://help.medium.com/hc/en-us?source=post_page-----f175e956a603---------------------------------------)\n\n[Status](https://status.medium.com/?source=post_page-----f175e956a603---------------------------------------)\n\n[About](https://medium.com/about?autoplay=1&source=post_page-----f175e956a603---------------------------------------)\n\n[Careers](https://medium.com/jobs-at-medium/work-at-medium-959d1a85284e?source=post_page-----f175e956a603---------------------------------------)\n\n[Press](mailto:pressinquiries@medium.com)\n\n[Blog](https://blog.medium.com/?source=post_page-----f175e956a603---------------------------------------)\n\n[Store](https://medium.com/store)\n\n[Privacy](https://policy.medium.com/medium-privacy-policy-f03bf92035c9?source=post_page-----f175e956a603---------------------------------------)\n\n[Rules](https://policy.medium.com/medium-rules-30e5502c4eb4?source=post_page-----f175e956a603---------------------------------------)\n\n[Terms](https://policy.medium.com/medium-terms-of-service-9db0094a1e0f?source=post_page-----f175e956a603---------------------------------------)\n\n[Text to speech](https://speechify.com/medium?source=post_page-----f175e956a603---------------------------------------)",
      "content_chars": 11544,
      "published_date": null
    },
    {
      "rank": 7,
      "url": "https://docs.olostep.com/get-started/welcome",
      "title": "Welcome to Olostep - Olostep Docs",
      "content": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/get-started/welcome#content-area)\n\nOlostep is infrastructure for the Web\u2019s second user \u2014 giving AI agents a way to search the web, extract structured data in real time, and build custom research agents.![Main dashboard interface](https://mintcdn.com/olostep-58/ndWPtiE100GTiImQ/images/explainer_v5.png?fit=max&auto=format&n=ndWPtiE100GTiImQ&q=85&s=ae92c6f02390b885a46f830090b0de34)\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#introduction)  Introduction\n\n- The **Olostep API** is the best **web search**, **scraping** and **crawling** API for AI used by some of the leading startups and scaleups in the world.\n- The **Olostep Agent** allows to automate **research workflows** in a no code way with just a prompt in natural language.\n\n**For AI agents:** fetch [docs.olostep.com/llms.txt](https://docs.olostep.com/llms.txt) for a complete index of this documentation before exploring further or [get started here](https://www.olostep.com/agent-onboarding/SKILL.md).\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#use-olostep-from-your-terminal-and-ai-agents)  Use Olostep from your terminal and AI agents\n\nBeyond the API, Olostep ships a CLI, an MCP server, and drop-in skills so any tool \u2014 Claude Code, Cursor, Windsurf, and more \u2014 can use the web natively.\n\n[**CLI \u2192** \\\\\n\\\\\n`npm i -g olostep-cli` \u2014 scrape, map, crawl, answer, and batch the web from your terminal. JSON output for scripts, CI, and agents.](https://docs.olostep.com/sdks/cli)\n\n[**MCP Server \u2192** \\\\\n\\\\\nGive any MCP client (Claude, Cursor, VS Code) live web tools. Hosted endpoint \u2014 no install.](https://docs.olostep.com/integrations/mcp-server)\n\n[**Skills \u2192** \\\\\n\\\\\nDrop-in skills that teach AI coding agents how and when to use Olostep. Install with `olostep add skills`.](https://docs.olostep.com/features/skills)\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#what-can-olostep-do)  What can Olostep do?\n\n[**Scrape** \\\\\n\\\\\nPull any URL as clean Markdown, HTML, screenshots, or structured JSON.](https://docs.olostep.com/get-started/welcome#scrape)\n\n[**Crawl** \\\\\n\\\\\nRecursively gather every page on a site, with filters and search.](https://docs.olostep.com/get-started/welcome#crawl)\n\n[**Answer** \\\\\n\\\\\nGet AI-synthesised answers from live web sources, with citations.](https://docs.olostep.com/get-started/welcome#answer)\n\n### [\u200b](https://docs.olostep.com/get-started/welcome\\#why-olostep)  Why Olostep?\n\n- **Built for AI**: Clean Markdown, structured JSON, citations \u2014 output your agents and apps consume directly.\n- **Reliable at scale**: Industry-leading success rate; handles JavaScript, anti-bot, and proxies under the hood.\n- **Fast**: Sub-second single scrape; up to 10,000 URLs in a single batch in 5\u20137 minutes.\n- **Cost-effective**: Significantly cheaper than alternatives at production scale.\n- **CLI + MCP + Skills**: Use Olostep from your terminal, scripts, or any MCP-aware agent \u2014 agent skills included.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#scrape)  Scrape\n\nPull any URL as clean Markdown. See the [Scrape feature docs](https://docs.olostep.com/features/scrapes) for all options.\n\nPython\n\nNode\n\ncURL\n\nCLI\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\nresult = client.scrapes.create(\n    url_to_scrape=\"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n    formats=[\"markdown\"],\n)\nprint(result.markdown_content)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst result = await client.scrapes.create({\n  url: 'https://en.wikipedia.org/wiki/Alexander_the_Great',\n  formats: ['markdown'],\n})\nconsole.log(result.markdown_content)\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/scrapes\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"url_to_scrape\": \"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n    \"formats\": [\"markdown\"]\n  }'\n```\n\n```\nolostep scrape \"https://en.wikipedia.org/wiki/Alexander_the_Great\"\n```\n\nResponse\n\n```\n{\n  \"id\": \"scrape_6h89o8u1kt\",\n  \"object\": \"scrape\",\n  \"result\": {\n    \"markdown_content\": \"## Alexander the Great...\",\n    \"markdown_hosted_url\": \"https://olostep-storage.s3.us-east-1.amazonaws.com/markDown_6h89o8u1kt.txt\",\n    \"page_metadata\": { \"status_code\": 200, \"title\": \"Alexander the Great - Wikipedia\" }\n  }\n}\n```\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#crawl)  Crawl\n\nRecursively gather every page on a site, with include/exclude filters and an optional `search_query` to focus the crawl. See [Crawl feature docs](https://docs.olostep.com/features/crawls).\n\nPython\n\nNode\n\ncURL\n\nCLI\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\ncrawl = client.crawls.create(\n    start_url=\"https://docs.olostep.com\",\n    max_pages=50,\n)\nfor page in crawl.pages():\n    print(page.url)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst crawl = await client.crawls.create({\n  url: 'https://docs.olostep.com',\n  maxPages: 50,\n})\nfor await (const page of crawl.pages()) {\n  console.log(page.url)\n}\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/crawls\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"start_url\": \"https://docs.olostep.com\",\n    \"max_pages\": 50\n  }'\n```\n\n```\nolostep crawl \"https://docs.olostep.com\" --max-pages 50\n```\n\nResponse\n\n```\n{\n  \"id\": \"crawl_abc123\",\n  \"object\": \"crawl\",\n  \"status\": \"completed\",\n  \"pages_count\": 47,\n  \"pages\": [\\\n    { \"url\": \"https://docs.olostep.com/get-started/welcome\", \"retrieve_id\": \"...\" },\\\n    { \"url\": \"https://docs.olostep.com/features/scrapes\", \"retrieve_id\": \"...\" }\\\n  ]\n}\n```\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#answer)  Answer\n\nAsk a question and get an AI-synthesised answer from live web sources, with citations. Pass a JSON schema to shape the output. See [Answers docs](https://docs.olostep.com/features/answers).\n\nPython\n\nNode\n\ncURL\n\nCLI\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\nanswer = client.answers.create(task=\"What does Olostep do?\")\nprint(answer.result)\nprint(answer.sources)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst answer = await client.answers.create({ task: 'What does Olostep do?' })\nconsole.log(answer.result)\nconsole.log(answer.sources)\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/answers\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"task\": \"What does Olostep do?\"}'\n```\n\n```\nolostep answer \"What does Olostep do?\"\n```\n\nResponse\n\n```\n{\n  \"id\": \"answer_abc123\",\n  \"object\": \"answer\",\n  \"task\": \"What does Olostep do?\",\n  \"result\": {\n    \"json_content\": \"{\\\"result\\\":\\\"Olostep is an API that lets AI agents search, scrape, and structure web data.\\\"}\",\n    \"sources\": [\\\n      \"https://docs.olostep.com/get-started/welcome\",\\\n      \"https://www.olostep.com/\"\\\n    ]\n  }\n}\n```\n\n* * *\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#more-capabilities)  More capabilities\n\n[**Batch** \\\\\n\\\\\nScrape up to 10,000 URLs in parallel; results back in 5\u20137 minutes.](https://docs.olostep.com/features/batches)\n\n[**Map** \\\\\n\\\\\nDiscover every URL on a site with include/exclude patterns.](https://docs.olostep.com/features/maps)\n\n[**Search** \\\\\n\\\\\nLive web search with structured links and optional inline scraping.](https://docs.olostep.com/features/search)\n\n[**Parsers** \\\\\n\\\\\nSelf-healing extractors that turn pages into typed JSON at scale.](https://docs.olostep.com/features/structured-content/parsers)\n\n[**Schedules** \\\\\n\\\\\nRun scrapes, crawls, and answers on a recurring schedule.](https://docs.olostep.com/features/schedules)\n\n[**Files** \\\\\n\\\\\nUpload files for batches or to connect your knowledge base.](https://docs.olostep.com/features/files)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#resources)  Resources\n\n[**Explore Features** \\\\\n\\\\\nCheck out all supported features for your scraping and AI search needs.](https://docs.olostep.com/features/scrapes)\n\n[**API Reference** \\\\\n\\\\\nStart using the API and test out the various params.](https://docs.olostep.com/api-reference/scrapes/create)\n\n[**Integrations** \\\\\n\\\\\nUse Olostep in n8n, make, relay, zapier, etc](https://docs.olostep.com/integrations/n8n)\n\n[**Examples** \\\\\n\\\\\nBrowse ready-to-use examples to get started quickly.](https://docs.olostep.com/examples/)\n\nWas this page helpful?\n\nYesNo",
      "content_chars": 8709,
      "published_date": null
    },
    {
      "rank": 8,
      "url": "https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats",
      "title": "Difference between Fragementation in percent vs Page count in Index ...",
      "content": "[Post reply](https://www.sqlservercentral.com/wp-login.php?redirect_to=https%3A%2F%2Fwww.sqlservercentral.com%2Fforums%2Ftopic%2Fdifference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats%23new-post)\n\n* * *\n\n# Difference between Fragementation in percent vs Page count in Index physical stats.\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 12:40 pm\n\n\n\n[#256845](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-256845)\n\n\n\n\n\n\n\nI would like to know about Fragementation in percent vs Page count. Whenever I tried to reorganize or rebuild the query, I see the following changes\n\n\n\nIndexTypeAvgPageFragmentationPageCounts\n\n\n\nCLUSTERED INDEX 66.6666666666667 3\n\n\n\nNONCLUSTERED INDEX 66.6666666666667 3\n\n\n\nHEAP63.7976346911958 12107\n\n\n\nCan someone explain me more detail way. If I further do any rebuild on clustered index, it remains same. Some times, the page count won't get reduced for example:\n\n\n\nIndexTypeAvgPageFragmentationPageCounts\n\n\n\nCLUSTERED INDEX 0.41958041958042 715\n\n\n\nWhat exactly this page count does and how can we reduce it or is it required to pay attention onto this page counts.\n\n- [SGT\\_squeequal](https://www.sqlservercentral.com/forums/user/sgtsqueequal)\n\n\n\n\n\n\n\n\n\nSSCertifiable\n\n\n\nPoints: 7167\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 1:29 pm\n\n\n\n[#1491343](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491343)\n\n\n\n\n\n\n\nHave a look at the stairways and indexing, there is plenty of information for you on there.\n\n\n\nas for pages, data is stored on pageses therefore if you have a page count of 10 then that is how many pages the data is held across\n\n\n\n\n\n\\*\\*\\*The first step is always the hardest \\*\\*\\*\\*\\*\\*\\*\n\n- [Gail Shaw](https://www.sqlservercentral.com/forums/user/gilamonster)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 1004485\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 1:56 pm\n\n\n\n[#1491361](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491361)\n\n\n\n\n\n\n\nPage count is the number of pages that the data in the table takes up, each page is 8kb. The only reliable way to decrease that is to delete data.\n\n\n\nDon't fuss over indexes with 3 pages, they won't defrag. The general recommendation is to worry about fragmentation once a table is over 1000 pages or so.\n\n\n\n\n\nGail Shaw\n\nMicrosoft Certified Master: SQL Server, MVP, M.Sc (Comp Sci)\n\n[**SQL In The Wild**](http://sqlinthewild.co.za/): Discussions on DB performance with occasional diversions into recoverability\n\n\n\n_We walk in the dark places no others will enter_\n\n_We stand on the bridge and no one may pass_\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:14 pm\n\n\n\n[#1491375](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491375)\n\n\n\n\n\n\n\nSo, if the table has more than 1000 pages, so how can we reduce it.\n\n- [Gail Shaw](https://www.sqlservercentral.com/forums/user/gilamonster)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 1004485\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:17 pm\n\n\n\n[#1491379](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491379)\n\n\n\n\n\n\n\nReduce what? Fragmentation or page count?\n\n\n\n\n\nGail Shaw\n\nMicrosoft Certified Master: SQL Server, MVP, M.Sc (Comp Sci)\n\n[**SQL In The Wild**](http://sqlinthewild.co.za/): Discussions on DB performance with occasional diversions into recoverability\n\n\n\n_We walk in the dark places no others will enter_\n\n_We stand on the bridge and no one may pass_\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:24 pm\n\n\n\n[#1491388](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491388)\n\n\n\n\n\n\n\nReduce page count\n\n- [Gail Shaw](https://www.sqlservercentral.com/forums/user/gilamonster)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 1004485\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:28 pm\n\n\n\n[#1491394](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491394)\n\n\n\n\n\n\n\nDelete data from the table. That's the only reliable way to reduce the size of the data in the table, which is what page count is.\n\n\n\nWhy are you fixated on page count? If a table has 8MB of data in it, it will have at least 1000 pages because 8MB of data, 8kb per page, 1000 pages.\n\n\n\n\n\nGail Shaw\n\nMicrosoft Certified Master: SQL Server, MVP, M.Sc (Comp Sci)\n\n[**SQL In The Wild**](http://sqlinthewild.co.za/): Discussions on DB performance with occasional diversions into recoverability\n\n\n\n_We walk in the dark places no others will enter_\n\n_We stand on the bridge and no one may pass_\n\n- [Lynn Pettis](https://www.sqlservercentral.com/forums/user/lynn-pettis)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 442462\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:28 pm\n\n\n\n[#1491395](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491395)\n\n\n\n\n\n\n\n\n\n\n\n\n\n> **DBA\\_SQL (5/22/2012)**\n>\n> * * *\n>\n> Reduce page count\n\n\n\n\n\n\n\n\n\nDelete data.\n\n\n\nTrying to understand why you are focusing on page count. As data is added to the database the page count is going to go up.\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 3:04 pm\n\n\n\n[#1491433](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491433)\n\n\n\n\n\n\n\nI think I am confusing. So, if we have more data, we get more pages...and vice versa right...So, at this point i think it is good to focus on only fragmentation part rather than pages, because it all depends on data.\n\n- [Lynn Pettis](https://www.sqlservercentral.com/forums/user/lynn-pettis)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 442462\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 3:06 pm\n\n\n\n[#1491434](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491434)\n\n\n\n\n\n\n\nYes. As Gail said earlier, don't even worry about fragmentation until you hit about 1000 pages in the table.\n\n- [Jared](https://www.sqlservercentral.com/forums/user/sqlknowitall)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 61793\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 3:06 pm\n\n\n\n[#1491436](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491436)\n\n\n\n\n\n\n\n\n\n\n\n\n\n> **DBA\\_SQL (5/22/2012)**\n>\n> * * *\n>\n> I think I am confusing. So, if we have more data, we get more pages...and vice versa right...So, at this point i think it is good to focus on only fragmentation part rather than pages, because it all depends on data.\n\n\n\n\n\n\n\n\n\nOk, but if the index is less than 1000 pages (or whatever you see fit), then don't worry about fragmentation.\n\n\n\n\n\nJared\n\nCE - Microsoft\n\n- [nmcquillen](https://www.sqlservercentral.com/forums/user/nmcquillen)\n\n\n\n\n\n\n\n\n\nGrasshopper\n\n\n\nPoints: 11\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nAugust 13, 2018 at 1:55 pm\n\n\n\n[#2001552](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-2001552)\n\n\n\n\n\n\n\nI realize this is a very old post, but a few things.\u00a0 If your index is getting scanned then ignore that 1000 pages business.\u00a0 I've seen scans on indexes choke with around a 300 page\\_count, so check your index scan stats.\u00a0 If you don't think you should be having scans, make sure whatever operation/query is sargable, no implicit\\_conversion, etc.\u00a0 Also if you have a higher page\\_count than you think should be required (considering offset/header) you might be dealing splitting while updating data bigger than the original slot so your page density isn't full when that occurs.\u00a0 Rebuild operations should handle this or you could go full tilt and take the db offline, defrag the mdf (to reduce physical fragmentation by hopefully getting contiguous mapping), rebuild indexes (sort in tempdb=on maxdop=1 for serial builds done on a separate disk than destination), and possibly mess around with fillfactor to mitigate splits.\n\n- [ScottPletcher](https://www.sqlservercentral.com/forums/user/scottpletcher)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 101249\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nAugust 13, 2018 at 2:02 pm\n\n\n\n[#2001553](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-2001553)\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n> **nmcquillen - Monday, August 13, 2018 1:55 PM**\n>\n> I realize this is a very old post, but a few things.\u00a0 If your index is getting scanned then ignore that 1000 pages business.\u00a0 I've seen scans on indexes choke with around a 300 page\\_count, so check your index scan stats.\u00a0 If you don't think you should be having scans, make sure whatever operation/query is sargable, no implicit\\_conversion, etc.\u00a0 Also if you have a higher page\\_count than you think should be required (considering offset/header) you might be dealing splitting while updating data bigger than the original slot so your page density isn't full when that occurs.\u00a0 Rebuild operations should handle this or you could go full tilt and take the db offline, defrag the mdf (to reduce physical fragmentation by hopefully getting contiguous mapping), rebuild indexes (sort in tempdb=on maxdop=1 for serial builds done on a separate disk than destination), and possibly mess around with fillfactor to mitigate splits.\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\nAbsolutely.\u00a0 Even the person who originally came up with the \"1000 pages\" admitted it was just a round number, with no real analysis to it.\u00a0 Determine the best clustering index and apply it, no matter how many rows the table (currently) has.\u00a0 You can also force a table(s) with less than 8 total pages to be put into a single extent, which for a busy, small table can reduce I/O.\n\n\n\nAs to ways to reduce data size, you have other options besides just deleting data (which usually isn't a viable business option at all);\n\n1) if you have an edition of SQL that supports it, compress the data.\n\n2) if you don't, encode data, esp. character data.\u00a0 That is, use a numeric id/code in place of a longer string value.\n\nThere are some others, but those should give you the biggest payback for the least overall effort.\n\n\n\n\n\n\n\nSQL DBA,SQL Server MVP(07, 08, 09) \"It's a dog-eat-dog world, and I'm wearing Milk-Bone underwear.\" \"Norm\", on \"Cheers\". Also from \"Cheers\", from \"Carla\": \"You need to know 3 things about Tortelli men: Tortelli men draw women like flies; Tortelli men treat women like flies; Tortelli men's brains are in their flies\".\n\n\nViewing 13 posts - 1 through 13 (of 13 total)\n\nYou must be logged in to reply to this topic. [Login to reply](https://www.sqlservercentral.com/wp-login.php?redirect_to=https%3A%2F%2Fwww.sqlservercentral.com%2Fforums%2Ftopic%2Fdifference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats%23new-post)",
      "content_chars": 13306,
      "published_date": null
    },
    {
      "rank": 9,
      "url": "https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17",
      "title": "Estimate the Size of a Clustered Index - SQL Server | Microsoft Learn",
      "content": "Table of contents Exit editor mode\n\nAsk LearnAsk Learn\n\nReading modeTable of contents[Read in English](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17)Add to CollectionsAdd to Plans[Edit](https://github.com/MicrosoftDocs/sql-docs/blob/live/docs/relational-databases/databases/estimate-the-size-of-a-clustered-index.md)\n\n* * *\n\nCopy MarkdownPrint\n\n* * *\n\nNote\n\nAccess to this page requires authorization. You can try [signing in](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#) or changing directories.\n\n\nAccess to this page requires authorization. You can try changing directories.\n\n\n# Estimate the size of a clustered index\n\nFeedback\n\nSummarize this article for me\n\n\n**Applies to:**![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [SQL Server](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [Azure SQL Database](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [Azure SQL Managed Instance](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [SQL database in Microsoft Fabric](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)\n\nYou can use the following steps to estimate the amount of space that is required to store data in a clustered index:\n\n1. Calculate the space used to store data in the leaf level of the clustered index.\n2. Calculate the space used to store index information for the clustered index.\n3. Total the calculated values.\n\n[Section titled: Step 1. Calculate the space used to store data in the leaf level](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#step-1-calculate-the-space-used-to-store-data-in-the-leaf-level)\n\n## Step 1. Calculate the space used to store data in the leaf level\n\n01. Specify the number of rows that are present in the table:\n\n    - _**Num\\_Rows**_ = number of rows in the table\n02. Specify the number of fixed-length and variable-length columns and calculate the space that is required for their storage:\n\n    Calculate the space that each of these groups of columns occupies within the data row. The size of a column depends on the data type and length specification.\n\n    - _**Num\\_Cols**_ = total number of columns (fixed-length and variable-length)\n    - _**Fixed\\_Data\\_Size**_ = total byte size of all fixed-length columns\n    - _**Num\\_Variable\\_Cols**_ = number of variable-length columns\n    - _**Max\\_Var\\_Size**_ = maximum byte size of all variable-length columns\n03. If the clustered index is nonunique, account for the _uniqueifier_ column:\n\n    The uniqueifier is a nullable, variable-length column. It's non-null and 4 bytes in size in rows that have nonunique key values. This value is part of the index key and is required to make sure that every row has a unique key value.\n\n\n    - _**Num\\_Cols**_ = _**Num\\_Cols**_ \\+ 1\n    - _**Num\\_Variable\\_Cols**_ = _**Num\\_Variable\\_Cols**_ \\+ 1\n    - _**Max\\_Var\\_Size**_ = _**Max\\_Var\\_Size**_ \\+ 4\n\nThese modifications assume that all values are nonunique.\n\n04. Part of the row, known as the null bitmap, is reserved to manage column nullability. Calculate its size:\n\n\n    - _**Null\\_Bitmap**_ = 2 + (( _**Num\\_Cols**_ \\+ 7) / 8)\n\nOnly the integer part of the previous expression should be used; discard any remainder.\n\n05. Calculate the variable-length data size:\n\n    If there are variable-length columns in the table, determine how much space is used to store the columns within the row:\n\n\n    - _**Variable\\_Data\\_Size**_ = 2 + ( _**Num\\_Variable\\_Cols**_ x 2) + _**Max\\_Var\\_Size**_\n\nThe bytes added to _**Max\\_Var\\_Size**_ are for tracking each variable column. This formula assumes that all variable-length columns are 100 percent full. If you anticipate that a smaller percentage of the variable-length column storage space will be used, you can adjust the _**Max\\_Var\\_Size**_ value by that percentage to yield a more accurate estimate of the overall table size.\n\nYou can combine **varchar**, **nvarchar**, **varbinary**, or **sql\\_variant** columns that cause the total defined table width to exceed 8,060 bytes. The length of each one of these columns must still fall within the limit of 8,000 bytes for a **varchar**, **varbinary**, or **sql\\_variant** column, and 4,000 bytes for **nvarchar** columns. However, their combined widths might exceed the 8,060-byte limit in a table.\n\nIf there are no variable-length columns, set _**Variable\\_Data\\_Size**_ to 0.\n\n06. Calculate the total row size:\n\n\n    - _**Row\\_Size**_ = _**Fixed\\_Data\\_Size**_ \\+ _**Variable\\_Data\\_Size**_ \\+ _**Null\\_Bitmap**_ \\+ 4\n\nThe value 4 is the row header overhead of a data row.\n\n07. Calculate the number of rows per page (8,096 free bytes per page):\n\n\n    - _**Rows\\_Per\\_Page**_ = 8096 / ( _**Row\\_Size**_ \\+ 2)\n\nBecause rows don't span pages, the number of rows per page should be rounded down to the nearest whole row. The value 2 in the formula is for the row's entry in the slot array of the page.\n\n08. Calculate the number of reserved free rows per page, based on the [fill factor](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/specify-fill-factor-for-an-index?view=sql-server-ver17) specified:\n\n\n    - _**Free\\_Rows\\_Per\\_Page**_ = 8096 x ((100 - _**Fill\\_Factor**_) / 100) / ( _**Row\\_Size**_ \\+ 2)\n\nThe fill factor used in the calculation is an integer value instead of a percentage. Because rows don't span pages, the number of rows per page should be rounded down to the nearest whole row. As the fill factor grows, more data is stored on each page and there are fewer pages. The value 2 in the formula is for the row's entry in the slot array of the page.\n\n09. Calculate the number of pages required to store all the rows:\n\n\n    - _**Num\\_Leaf\\_Pages**_ = _**Num\\_Rows**_ / ( _**Rows\\_Per\\_Page**_ \\- _**Free\\_Rows\\_Per\\_Page**_)\n\nThe number of pages estimated should be rounded up to the nearest whole page.\n\n10. Calculate the amount of space that is required to store the data in the leaf level (8,192 total bytes per page):\n\n    - _**Leaf\\_space\\_used**_ = 8192 x _**Num\\_Leaf\\_Pages**_\n\n[Section titled: Step 2. Calculate the space used to store index information](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#step-2-calculate-the-space-used-to-store-index-information)\n\n## Step 2. Calculate the space used to store index information\n\nYou can use the following steps to estimate the amount of space that is required to store the upper levels of the index:\n\n1. Specify the number of fixed-length and variable-length columns in the index key and calculate the space that is required for their storage:\n\nThe key columns of an index can include fixed-length and variable-length columns. To estimate the interior level index row size, calculate the space that each of these groups of columns occupies within the index row. The size of a column depends on the data type and length specification.\n\n   - _**Num\\_Key\\_Cols**_ = total number of key columns (fixed-length and variable-length)\n   - _**Fixed\\_Key\\_Size**_ = total byte size of all fixed-length key columns\n   - _**Num\\_Variable\\_Key\\_Cols**_ = number of variable-length key columns\n   - _**Max\\_Var\\_Key\\_Size**_ = maximum byte size of all variable-length key columns\n2. Account for any uniqueifier needed if the index is nonunique:\n\nThe uniqueifier is a nullable, variable-length column. It's non-null and 4 bytes in size in rows that have nonunique index key values. This value is part of the index key and is required to make sure that every row has a unique key value.\n\n\n   - _**Num\\_Key\\_Cols**_ = _**Num\\_Key\\_Cols**_ \\+ 1\n   - _**Num\\_Variable\\_Key\\_Cols**_ = _**Num\\_Variable\\_Key\\_Cols**_ \\+ 1\n   - _**Max\\_Var\\_Key\\_Size**_ = _**Max\\_Var\\_Key\\_Size**_ \\+ 4\n\nThese modifications assume that all values are nonunique.\n\n3. Calculate the null bitmap size:\n\nIf there are nullable columns in the index key, part of the index row is reserved for the null bitmap. Calculate its size:\n\n\n   - _**Index\\_Null\\_Bitmap**_ = 2 + ((number of columns in the index row + 7) / 8)\n\nOnly the integer part of the previous expression should be used. Discard any remainder.\n\nIf there are no nullable key columns, set _**Index\\_Null\\_Bitmap**_ to 0.\n\n4. Calculate the variable-length data size:\n\nIf there are variable-length columns in the index, determine how much space is used to store the columns within the index row:\n\n\n   - _**Variable\\_Key\\_Size**_ = 2 + ( _**Num\\_Variable\\_Key\\_Cols**_ x 2) + _**Max\\_Var\\_Key\\_Size**_\n\nThe bytes added to _**Max\\_Var\\_Key\\_Size**_ are for tracking each variable-length column. This formula assumes that all variable-length columns are 100 percent full. If you anticipate that a smaller percentage of the variable-length column storage space will be used, you can adjust the _**Max\\_Var\\_Key\\_Size**_ value by that percentage to yield a more accurate estimate of the overall table size.\n\nIf there are no variable-length columns, set _**Variable\\_Key\\_Size**_ to 0.\n\n5. Calculate the index row size:\n\n   - _**Index\\_Row\\_Size**_ = _**Fixed\\_Key\\_Size**_ \\+ _**Variable\\_Key\\_Size**_ \\+ _**Index\\_Null\\_Bitmap**_ \\+ 1 (for row header overhead of an index row) + 6 (for the child page ID pointer)\n6. Calculate the number of index rows per page (8,096 free bytes per page):\n\n\n   - _**Index\\_Rows\\_Per\\_Page**_ = 8096 / ( _**Index\\_Row\\_Size**_ \\+ 2)\n\nBecause index rows don't span pages, the number of index rows per page should be rounded down to the nearest whole row. The `2` in the formula is for the row's entry in the page's slot array.\n\n7. Calculate the number of levels in the index:\n\n\n   - _**Non-leaf\\_Levels**_ = 1 + log (Index\\_Rows\\_Per\\_Page) ( _**Num\\_Leaf\\_Pages**_ / _**Index\\_Rows\\_Per\\_Page**_)\n\nRound this value up to the nearest whole number. This value doesn't include the leaf level of the clustered index.\n\n8. Calculate the number of nonleaf pages in the index:\n\n\n   - _**Num\\_Index\\_Pages =**_ \u2211Level ( _**Num\\_Leaf\\_Pages**_ / ( _**Index\\_Rows\\_Per\\_Page**_ ^ _**Level**_))\n\n     where 1 <= Level <= _**Non-leaf\\_Levels**_\n\n\nRound each summand up to the nearest whole number. As a simple example, consider an index where _**Num\\_Leaf\\_Pages**_ = 1000 and _**Index\\_Rows\\_Per\\_Page**_ = 25\\. The first index level above the leaf level stores 1,000 index rows, which is one index row per leaf page, and 25 index rows can fit per page. This means that 40 pages are required to store those 1,000 index rows. The next level of the index has to store 40 rows. This means it requires two pages. The final level of the index has to store two rows. This means it requires one page. This gives 43 nonleaf index pages. When these numbers are used in the previous formulas, the outcome is as follows:\n\n   - _**Non-leaf\\_Levels**_ = 1 + log(25) (1000 / 25) = 3\n\n   - _**Num\\_Index\\_Pages**_ = 1000/(25^3)+ 1000/(25^2) + 1000/(25^1) = 1 + 2 + 40 = 43, which is the number of pages described in the example.\n9. Calculate the size of the index (8,192 total bytes per page):\n\n   - _**Index\\_Space\\_Used**_ = 8192 x _**Num\\_Index\\_Pages**_\n\n[Section titled: Step 3. Total the calculated values](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#step-3-total-the-calculated-values)\n\n## Step 3. Total the calculated values\n\nTotal the values obtained from the previous two steps:\n\n- Clustered index size (bytes) = _**Leaf\\_Space\\_Used**_ \\+ _**Index\\_Space\\_used**_\n\nThis calculation doesn't consider the following conditions:\n\n- **Partitioning**: The space overhead from partitioning is minimal, but complex to calculate. It isn't important to include.\n\n- **Allocation pages**: There's at least one IAM page used to track the pages allocated to a heap. The space overhead is minimal, and there's no algorithm to deterministically calculate exactly how many IAM pages will be used.\n\n- **Large object (LOB) values**: The algorithm to determine exactly how much space will be used to store the LOB data types **varchar(max)**, **varbinary(max)**, **nvarchar(max)**, **text**, **ntext**, **xml**, and **image** values is complex. It's sufficient to just add the average size of the LOB values that are expected, multiply by _**Num\\_Rows**_, and add that to the total clustered index size.\n\n- **Compression**: You can't precalculate the size of a compressed index.\n\n- **Sparse columns**: For information about the space requirements of sparse columns, see [Use sparse columns](https://learn.microsoft.com/en-us/sql/relational-databases/tables/use-sparse-columns?view=sql-server-ver17).\n\n\n[Section titled: Related content](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#related-content)\n\n## Related content\n\n- [Clustered and nonclustered indexes](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/clustered-and-nonclustered-indexes-described?view=sql-server-ver17)\n- [Estimate the size of a table](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-table?view=sql-server-ver17)\n- [Create a clustered index](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/create-clustered-indexes?view=sql-server-ver17)\n- [Create nonclustered indexes](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/create-nonclustered-indexes?view=sql-server-ver17)\n- [Estimate the size of a nonclustered index](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-nonclustered-index?view=sql-server-ver17)\n- [Estimate the size of a heap](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-heap?view=sql-server-ver17)\n- [Estimate the size of a database](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-database?view=sql-server-ver17)\n\nReading mode disabled\n\n* * *\n\n## Feedback\n\nWas this page helpful?\n\n\nYesNoNo\n\nNeed help with this topic?\n\n\nWant to try using Ask Learn to clarify or guide you through this topic?\n\n\nAsk LearnAsk Learn\n\nSuggest a fix?\n\n* * *\n\n## Additional resources\n\n* * *\n\n- Last updated on 07/20/2026\n\nAsk Learn is an AI assistant that can answer questions, clarify concepts, and define terms using trusted Microsoft documentation.\n\nPlease sign in to use Ask Learn.\n\n[Sign in](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#)",
      "content_chars": 15181,
      "published_date": null
    },
    {
      "rank": 10,
      "url": "https://docs.olostep.com/api-reference/scrapes/create",
      "title": "Create Scrape - Olostep Docs",
      "content": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/api-reference/scrapes/create#content-area)\n\nInitiate a web page scrape\n\ncURL\n\n```\ncurl --request POST \\\n  --url https://api.olostep.com/v1/scrapes \\\n  --header 'Authorization: Bearer <token>' \\\n  --header 'Content-Type: application/json' \\\n  --data '\n{\n  \"url_to_scrape\": \"<string>\",\n  \"wait_before_scraping\": 123,\n  \"formats\": [],\n  \"actions\": [\\\n    {\\\n      \"type\": \"wait\",\\\n      \"milliseconds\": 1\\\n    }\\\n  ],\n  \"country\": \"<string>\",\n  \"remove_images\": false,\n  \"remove_class_names\": [\\\n    \"<string>\"\\\n  ],\n  \"llm_extract\": {\n    \"schema\": {}\n  },\n  \"links_on_page\": {\n    \"query_to_order_links_by\": \"<string>\",\n    \"include_links\": [\\\n      \"<string>\"\\\n    ],\n    \"exclude_links\": [\\\n      \"<string>\"\\\n    ]\n  },\n  \"screen_size\": {\n    \"screen_width\": 123,\n    \"screen_height\": 123\n  },\n  \"screenshot\": {\n    \"full_page\": true\n  },\n  \"metadata\": {},\n  \"max_age\": 0\n}\n'\n```\n\n```\nimport requests\n\nurl = \"https://api.olostep.com/v1/scrapes\"\n\npayload = {\n    \"url_to_scrape\": \"<string>\",\n    \"wait_before_scraping\": 123,\n    \"formats\": [],\n    \"actions\": [\\\n        {\\\n            \"type\": \"wait\",\\\n            \"milliseconds\": 1\\\n        }\\\n    ],\n    \"country\": \"<string>\",\n    \"remove_images\": False,\n    \"remove_class_names\": [\"<string>\"],\n    \"llm_extract\": { \"schema\": {} },\n    \"links_on_page\": {\n        \"query_to_order_links_by\": \"<string>\",\n        \"include_links\": [\"<string>\"],\n        \"exclude_links\": [\"<string>\"]\n    },\n    \"screen_size\": {\n        \"screen_width\": 123,\n        \"screen_height\": 123\n    },\n    \"screenshot\": { \"full_page\": True },\n    \"metadata\": {},\n    \"max_age\": 0\n}\nheaders = {\n    \"Authorization\": \"Bearer <token>\",\n    \"Content-Type\": \"application/json\"\n}\n\nresponse = requests.post(url, json=payload, headers=headers)\n\nprint(response.text)\n```\n\n```\nconst options = {\n  method: 'POST',\n  headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},\n  body: JSON.stringify({\n    url_to_scrape: '<string>',\n    wait_before_scraping: 123,\n    formats: [],\n    actions: [{type: 'wait', milliseconds: 1}],\n    country: '<string>',\n    remove_images: false,\n    remove_class_names: ['<string>'],\n    llm_extract: {schema: {}},\n    links_on_page: {\n      query_to_order_links_by: '<string>',\n      include_links: ['<string>'],\n      exclude_links: ['<string>']\n    },\n    screen_size: {screen_width: 123, screen_height: 123},\n    screenshot: {full_page: true},\n    metadata: {},\n    max_age: 0\n  })\n};\n\nfetch('https://api.olostep.com/v1/scrapes', options)\n  .then(res => res.json())\n  .then(res => console.log(res))\n  .catch(err => console.error(err));\n```\n\n```\n<?php\n\n$curl = curl_init();\n\ncurl_setopt_array($curl, [\\\n  CURLOPT_URL => \"https://api.olostep.com/v1/scrapes\",\\\n  CURLOPT_RETURNTRANSFER => true,\\\n  CURLOPT_ENCODING => \"\",\\\n  CURLOPT_MAXREDIRS => 10,\\\n  CURLOPT_TIMEOUT => 30,\\\n  CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,\\\n  CURLOPT_CUSTOMREQUEST => \"POST\",\\\n  CURLOPT_POSTFIELDS => json_encode([\\\n    'url_to_scrape' => '<string>',\\\n    'wait_before_scraping' => 123,\\\n    'formats' => [\\\n\\\n    ],\\\n    'actions' => [\\\n        [\\\n                'type' => 'wait',\\\n                'milliseconds' => 1\\\n        ]\\\n    ],\\\n    'country' => '<string>',\\\n    'remove_images' => false,\\\n    'remove_class_names' => [\\\n        '<string>'\\\n    ],\\\n    'llm_extract' => [\\\n        'schema' => [\\\n\\\n        ]\\\n    ],\\\n    'links_on_page' => [\\\n        'query_to_order_links_by' => '<string>',\\\n        'include_links' => [\\\n                '<string>'\\\n        ],\\\n        'exclude_links' => [\\\n                '<string>'\\\n        ]\\\n    ],\\\n    'screen_size' => [\\\n        'screen_width' => 123,\\\n        'screen_height' => 123\\\n    ],\\\n    'screenshot' => [\\\n        'full_page' => true\\\n    ],\\\n    'metadata' => [\\\n\\\n    ],\\\n    'max_age' => 0\\\n  ]),\\\n  CURLOPT_HTTPHEADER => [\\\n    \"Authorization: Bearer <token>\",\\\n    \"Content-Type: application/json\"\\\n  ],\\\n]);\n\n$response = curl_exec($curl);\n$err = curl_error($curl);\n\ncurl_close($curl);\n\nif ($err) {\n  echo \"cURL Error #:\" . $err;\n} else {\n  echo $response;\n}\n```\n\n```\npackage main\n\nimport (\n\t\"fmt\"\n\t\"strings\"\n\t\"net/http\"\n\t\"io\"\n)\n\nfunc main() {\n\n\turl := \"https://api.olostep.com/v1/scrapes\"\n\n\tpayload := strings.NewReader(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n\n\treq, _ := http.NewRequest(\"POST\", url, payload)\n\n\treq.Header.Add(\"Authorization\", \"Bearer <token>\")\n\treq.Header.Add(\"Content-Type\", \"application/json\")\n\n\tres, _ := http.DefaultClient.Do(req)\n\n\tdefer res.Body.Close()\n\tbody, _ := io.ReadAll(res.Body)\n\n\tfmt.Println(string(body))\n\n}\n```\n\n```\nHttpResponse<String> response = Unirest.post(\"https://api.olostep.com/v1/scrapes\")\n  .header(\"Authorization\", \"Bearer <token>\")\n  .header(\"Content-Type\", \"application/json\")\n  .body(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n  .asString();\n```\n\n```\nrequire 'uri'\nrequire 'net/http'\n\nurl = URI(\"https://api.olostep.com/v1/scrapes\")\n\nhttp = Net::HTTP.new(url.host, url.port)\nhttp.use_ssl = true\n\nrequest = Net::HTTP::Post.new(url)\nrequest[\"Authorization\"] = 'Bearer <token>'\nrequest[\"Content-Type\"] = 'application/json'\nrequest.body = \"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\"\n\nresponse = http.request(request)\nputs response.read_body\n```\n\n200\n\n400\n\n502\n\n504\n\n```\n{\n  \"id\": \"<string>\",\n  \"object\": \"<string>\",\n  \"created\": 123,\n  \"metadata\": {},\n  \"url_to_scrape\": \"<string>\",\n  \"result\": {\n    \"html_content\": \"<string>\",\n    \"markdown_content\": \"<string>\",\n    \"text_content\": \"<string>\",\n    \"json_content\": \"<string>\",\n    \"screenshot_hosted_url\": \"<string>\",\n    \"html_hosted_url\": \"<string>\",\n    \"markdown_hosted_url\": \"<string>\",\n    \"text_hosted_url\": \"<string>\",\n    \"links_on_page\": [\\\n      \"<string>\"\\\n    ],\n    \"page_metadata\": {\n      \"status_code\": 123,\n      \"title\": \"<string>\"\n    }\n  },\n  \"credits_consumed\": 123,\n  \"cost_usd\": 123\n}\n```\n\n```\n{\n  \"id\": \"error_x2nmu5bqn6\",\n  \"object\": \"error\",\n  \"created\": 1777923912,\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"dns_resolution_failed\",\n    \"message\": \"The URL contains a typo, or the domain does not exist.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_ogeb6rik8c\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"tls_error\",\n    \"detail\": \"err_ssl_tlsv1_alert_internal_error\",\n    \"message\": \"The website closed or rejected the TLS handshake. The server may be misconfigured or use an unsupported SSL/TLS version.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_qat3d1amjt\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"request_timeout\",\n    \"code\": \"scrape_poll_timeout\",\n    \"message\": \"Request timed out while waiting for scrape result. The page may be slow, blocked for our fetchers, or temporarily unavailable.\"\n  }\n}\n```\n\nPOST\n\n/\n\nv1\n\n/\n\nscrapes\n\nTry it\n\nInitiate a web page scrape\n\ncURL\n\n```\ncurl --request POST \\\n  --url https://api.olostep.com/v1/scrapes \\\n  --header 'Authorization: Bearer <token>' \\\n  --header 'Content-Type: application/json' \\\n  --data '\n{\n  \"url_to_scrape\": \"<string>\",\n  \"wait_before_scraping\": 123,\n  \"formats\": [],\n  \"actions\": [\\\n    {\\\n      \"type\": \"wait\",\\\n      \"milliseconds\": 1\\\n    }\\\n  ],\n  \"country\": \"<string>\",\n  \"remove_images\": false,\n  \"remove_class_names\": [\\\n    \"<string>\"\\\n  ],\n  \"llm_extract\": {\n    \"schema\": {}\n  },\n  \"links_on_page\": {\n    \"query_to_order_links_by\": \"<string>\",\n    \"include_links\": [\\\n      \"<string>\"\\\n    ],\n    \"exclude_links\": [\\\n      \"<string>\"\\\n    ]\n  },\n  \"screen_size\": {\n    \"screen_width\": 123,\n    \"screen_height\": 123\n  },\n  \"screenshot\": {\n    \"full_page\": true\n  },\n  \"metadata\": {},\n  \"max_age\": 0\n}\n'\n```\n\n```\nimport requests\n\nurl = \"https://api.olostep.com/v1/scrapes\"\n\npayload = {\n    \"url_to_scrape\": \"<string>\",\n    \"wait_before_scraping\": 123,\n    \"formats\": [],\n    \"actions\": [\\\n        {\\\n            \"type\": \"wait\",\\\n            \"milliseconds\": 1\\\n        }\\\n    ],\n    \"country\": \"<string>\",\n    \"remove_images\": False,\n    \"remove_class_names\": [\"<string>\"],\n    \"llm_extract\": { \"schema\": {} },\n    \"links_on_page\": {\n        \"query_to_order_links_by\": \"<string>\",\n        \"include_links\": [\"<string>\"],\n        \"exclude_links\": [\"<string>\"]\n    },\n    \"screen_size\": {\n        \"screen_width\": 123,\n        \"screen_height\": 123\n    },\n    \"screenshot\": { \"full_page\": True },\n    \"metadata\": {},\n    \"max_age\": 0\n}\nheaders = {\n    \"Authorization\": \"Bearer <token>\",\n    \"Content-Type\": \"application/json\"\n}\n\nresponse = requests.post(url, json=payload, headers=headers)\n\nprint(response.text)\n```\n\n```\nconst options = {\n  method: 'POST',\n  headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},\n  body: JSON.stringify({\n    url_to_scrape: '<string>',\n    wait_before_scraping: 123,\n    formats: [],\n    actions: [{type: 'wait', milliseconds: 1}],\n    country: '<string>',\n    remove_images: false,\n    remove_class_names: ['<string>'],\n    llm_extract: {schema: {}},\n    links_on_page: {\n      query_to_order_links_by: '<string>',\n      include_links: ['<string>'],\n      exclude_links: ['<string>']\n    },\n    screen_size: {screen_width: 123, screen_height: 123},\n    screenshot: {full_page: true},\n    metadata: {},\n    max_age: 0\n  })\n};\n\nfetch('https://api.olostep.com/v1/scrapes', options)\n  .then(res => res.json())\n  .then(res => console.log(res))\n  .catch(err => console.error(err));\n```\n\n```\n<?php\n\n$curl = curl_init();\n\ncurl_setopt_array($curl, [\\\n  CURLOPT_URL => \"https://api.olostep.com/v1/scrapes\",\\\n  CURLOPT_RETURNTRANSFER => true,\\\n  CURLOPT_ENCODING => \"\",\\\n  CURLOPT_MAXREDIRS => 10,\\\n  CURLOPT_TIMEOUT => 30,\\\n  CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,\\\n  CURLOPT_CUSTOMREQUEST => \"POST\",\\\n  CURLOPT_POSTFIELDS => json_encode([\\\n    'url_to_scrape' => '<string>',\\\n    'wait_before_scraping' => 123,\\\n    'formats' => [\\\n\\\n    ],\\\n    'actions' => [\\\n        [\\\n                'type' => 'wait',\\\n                'milliseconds' => 1\\\n        ]\\\n    ],\\\n    'country' => '<string>',\\\n    'remove_images' => false,\\\n    'remove_class_names' => [\\\n        '<string>'\\\n    ],\\\n    'llm_extract' => [\\\n        'schema' => [\\\n\\\n        ]\\\n    ],\\\n    'links_on_page' => [\\\n        'query_to_order_links_by' => '<string>',\\\n        'include_links' => [\\\n                '<string>'\\\n        ],\\\n        'exclude_links' => [\\\n                '<string>'\\\n        ]\\\n    ],\\\n    'screen_size' => [\\\n        'screen_width' => 123,\\\n        'screen_height' => 123\\\n    ],\\\n    'screenshot' => [\\\n        'full_page' => true\\\n    ],\\\n    'metadata' => [\\\n\\\n    ],\\\n    'max_age' => 0\\\n  ]),\\\n  CURLOPT_HTTPHEADER => [\\\n    \"Authorization: Bearer <token>\",\\\n    \"Content-Type: application/json\"\\\n  ],\\\n]);\n\n$response = curl_exec($curl);\n$err = curl_error($curl);\n\ncurl_close($curl);\n\nif ($err) {\n  echo \"cURL Error #:\" . $err;\n} else {\n  echo $response;\n}\n```\n\n```\npackage main\n\nimport (\n\t\"fmt\"\n\t\"strings\"\n\t\"net/http\"\n\t\"io\"\n)\n\nfunc main() {\n\n\turl := \"https://api.olostep.com/v1/scrapes\"\n\n\tpayload := strings.NewReader(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n\n\treq, _ := http.NewRequest(\"POST\", url, payload)\n\n\treq.Header.Add(\"Authorization\", \"Bearer <token>\")\n\treq.Header.Add(\"Content-Type\", \"application/json\")\n\n\tres, _ := http.DefaultClient.Do(req)\n\n\tdefer res.Body.Close()\n\tbody, _ := io.ReadAll(res.Body)\n\n\tfmt.Println(string(body))\n\n}\n```\n\n```\nHttpResponse<String> response = Unirest.post(\"https://api.olostep.com/v1/scrapes\")\n  .header(\"Authorization\", \"Bearer <token>\")\n  .header(\"Content-Type\", \"application/json\")\n  .body(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n  .asString();\n```\n\n```\nrequire 'uri'\nrequire 'net/http'\n\nurl = URI(\"https://api.olostep.com/v1/scrapes\")\n\nhttp = Net::HTTP.new(url.host, url.port)\nhttp.use_ssl = true\n\nrequest = Net::HTTP::Post.new(url)\nrequest[\"Authorization\"] = 'Bearer <token>'\nrequest[\"Content-Type\"] = 'application/json'\nrequest.body = \"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\"\n\nresponse = http.request(request)\nputs response.read_body\n```\n\n200\n\n400\n\n502\n\n504\n\n```\n{\n  \"id\": \"<string>\",\n  \"object\": \"<string>\",\n  \"created\": 123,\n  \"metadata\": {},\n  \"url_to_scrape\": \"<string>\",\n  \"result\": {\n    \"html_content\": \"<string>\",\n    \"markdown_content\": \"<string>\",\n    \"text_content\": \"<string>\",\n    \"json_content\": \"<string>\",\n    \"screenshot_hosted_url\": \"<string>\",\n    \"html_hosted_url\": \"<string>\",\n    \"markdown_hosted_url\": \"<string>\",\n    \"text_hosted_url\": \"<string>\",\n    \"links_on_page\": [\\\n      \"<string>\"\\\n    ],\n    \"page_metadata\": {\n      \"status_code\": 123,\n      \"title\": \"<string>\"\n    }\n  },\n  \"credits_consumed\": 123,\n  \"cost_usd\": 123\n}\n```\n\n```\n{\n  \"id\": \"error_x2nmu5bqn6\",\n  \"object\": \"error\",\n  \"created\": 1777923912,\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"dns_resolution_failed\",\n    \"message\": \"The URL contains a typo, or the domain does not exist.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_ogeb6rik8c\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"tls_error\",\n    \"detail\": \"err_ssl_tlsv1_alert_internal_error\",\n    \"message\": \"The website closed or rejected the TLS handshake. The server may be misconfigured or use an unsupported SSL/TLS version.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_qat3d1amjt\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"request_timeout\",\n    \"code\": \"scrape_poll_timeout\",\n    \"message\": \"Request timed out while waiting for scrape result. The page may be slow, blocked for our fetchers, or temporarily unavailable.\"\n  }\n}\n```\n\n**Optional caching:** Pass `max_age` (in seconds) to reuse a recent scrape with the same parameters instead of fetching the page again. Defaults to `0` (always fresh). In the dashboard playground, the default is 24 hours. See [Caching](https://docs.olostep.com/features/scrapes#caching) for details.\n\n#### Authorizations\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#authorization-authorization)\n\nAuthorization\n\nstring\n\nheader\n\nrequired\n\nBearer authentication header of the form Bearer , where  is your auth token.\n\n#### Body\n\napplication/json\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-url-to-scrape)\n\nurl\\_to\\_scrape\n\nstring<uri>\n\nrequired\n\nThe URL to start scraping from.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-wait-before-scraping)\n\nwait\\_before\\_scraping\n\ninteger\n\nTime to wait in milliseconds before starting the scraping.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-formats)\n\nformats\n\nenum<string>\\[\\]\n\nFormats in which you want the content.\n\nAvailable options:\n\n`html`,\n\n`markdown`,\n\n`text`,\n\n`json`,\n\n`raw_pdf`,\n\n`screenshot`\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-remove-css-selectors)\n\nremove\\_css\\_selectors\n\nenum<string>\n\nOption to remove certain CSS selectors from the content. Optionally, you can also pass a JSON stringified array of specific selectors you want to remove. The CSS selectors removed when this option is set to default are \\['nav','footer','script','style','noscript','svg',\\[role=alert\\],\\[role=banner\\],\\[role=dialog\\],\\[role=alertdialog\\],\\[role=region\\]\\[aria-label\\*=skip i\\],\\[aria-modal=true\\]\\]\n\nAvailable options:\n\n`default`,\n\n`none`,\n\n`array`\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-actions)\n\nactions\n\n(Wait \u00b7 object \\| Click \u00b7 object \\| Fill Input \u00b7 object \\| Scroll \u00b7 object)\\[\\]\n\nActions to perform on the page before getting the content.\n\n- Wait\n\n- Click\n\n- Fill Input\n\n- Scroll\n\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-country)\n\ncountry\n\nstring\n\nResidential country to load the request from.\n\nSupported values are:\n\n- US (United States)\n- CA (Canada)\n- IT (Italy)\n- IN (India)\n- GB (England)\n- JP (Japan)\n- MX (Mexico)\n- AU (Australia)\n- ID (Indonesia)\n- UA (UAE)\n- RU (Russia)\n- RANDOM\n\nSome operations, like scraping Google Search and Google News, support all countries.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-transformer)\n\ntransformer\n\nenum<string>\n\nSpecify the HTML transformer to use, if any. Postlight's Mercury Parser library is used to remove ads and other unwanted content from the scraped content.\n\nAvailable options:\n\n`postlight`,\n\n`none`\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-remove-images)\n\nremove\\_images\n\nboolean\n\ndefault:false\n\nOption to remove images from the scraped content. Defaults to false.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-remove-class-names)\n\nremove\\_class\\_names\n\nstring\\[\\]\n\nList of class names to remove from the content.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-parser)\n\nparser\n\nobject\n\nWhen defining json as a format, you can use this parameter to specify the parser to use. Parsers are useful to extract structured content from web pages. Olostep has a few parsers built in for most common web pages, and you can also create your own parsers.\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-llm-extract)\n\nllm\\_extract\n\nobject\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-links-on-page)\n\nlinks\\_on\\_page\n\nobject\n\nWith this option, you can get all the links present on the page you scrape. Links are always returned as absolute URLs.\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-screen-size)\n\nscreen\\_size\n\nobject\n\nConfiguration for screen size. Preset dimensions are available through screen\\_type: desktop (1920x1080), mobile (414x896), or default (768x1024).\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-screenshot)\n\nscreenshot\n\nobject\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-metadata)\n\nmetadata\n\nobject\n\nUser-defined metadata. Not supported yet\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-max-age)\n\nmax\\_age\n\ninteger\n\ndefault:0\n\nMaximum acceptable age of cached content, in seconds. When a matching scrape already exists and is newer than max\\_age seconds, Olostep returns the stored result instead of launching a new browser scrape. Defaults to 0 (always scrape fresh). In the dashboard playground, the default is 86400 (24 hours). The maximum allowed value is 604800 (7 days). See the Caching section in the Scrapes feature docs for details.\n\nRequired range: `x >= 0`\n\n#### Response\n\n200\n\napplication/json\n\nSuccessful response with the scrape initiation details.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-id)\n\nid\n\nstring\n\nScrape ID\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-object)\n\nobject\n\nstring\n\nThe kind of object. \"scrape\" for this endpoint.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-created)\n\ncreated\n\nnumber\n\nCreated epoch\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-metadata)\n\nmetadata\n\nobject\n\nUser-defined metadata.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-url-to-scrape)\n\nurl\\_to\\_scrape\n\nstring\n\nThe URL that was scraped.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-result)\n\nresult\n\nobject\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-credits-consumed-one-of-0)\n\ncredits\\_consumed\n\ninteger \\| null\n\nNumber of credits consumed by this request. Populated after execution completes. Credits are the source of truth for billing.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-cost-usd-one-of-0)\n\ncost\\_usd\n\nnumber \\| null\n\nEstimated cost in USD for this request. Populated after execution completes. Calculated from credits consumed and your plan rate \u2014 99% accurate, but credits\\_consumed is the authoritative value.\n\nWas this page helpful?\n\nYesNo",
      "content_chars": 24236,
      "published_date": null
    }
  ],
  "answer_text": null,
  "citations": [],
  "raw_response": {
    "success": true,
    "data": {
      "web": [
        {
          "url": "https://www.olostep.com/",
          "title": "Olostep: Web Data Infrastructure for AI Agents",
          "description": "# Web Data Infrastructure for AI\n## Built for Developers\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6crawl = client.crawls.create(\n7    start_url=\"https://olostep.com\",\n8    max_pages=100,\n9    include_urls=[\"/**\"],\n10    exclude_urls=[\"/collections/**\"],\n11    include_external=False,\n12)\n13\n14print(crawl.id, crawl.status)\n15\n16# Wait for completion and iterate pages\n17for page in crawl.pages():\n18    print(page.url)\n19    content = page.retrieve([\"markdown\"])\n20    print(content.markdown_content[:200])\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const crawl = await client.crawls.create({\n7  url: 'https://olostep.com',\n8  maxPages: 100,\n9  includeUrls: ['/**'],\n10  excludeUrls: ['/collections/**'],\n11  includeExternal: false,\n12})\n13\n14console.log(crawl.id, crawl.status)\n15\n16// Wait for completion and iterate pages\n17for await (const page of crawl.pages()) {\n18  console.log(page.url)\n19  const content = await client.retrieve({ retrieveId: page.retrieve_id, formats: ['markdown'] })\n20  console.log(content.markdown_content.slice(0, 200))\n21}\n```\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6sitemap = client.maps.create(\n7    url=\"https://docs.olostep.com\",\n8    include_urls=[\"/features/**\"],\n9    top_n=100,\n10)\n11\n12print(f\"Map ID: {sitemap.id}\")\n13\n14# Iterate all URLs (handles pagination automatically)\n15for url in sitemap.urls():\n16    print(url)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const map = await client.maps.create({\n7  url: 'https://docs.olostep.com',\n8  includeUrls: ['/features/**'],\n9  topN: 100,\n10})\n11\n12console.log(`Map ID: ${map.id}`)\n13\n14// Iterate all URLs (handles pagination automatically)\n15for await (const url of map.urls()) {\n16  console.log(url)\n17}\n```\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6search = client.searches.create(\"Latest updates with SpaceX\")\n7\n8print(search.id, len(search.links))\n```",
          "position": 1,
          "markdown": "Click to try\n\nWait...\n\n![](https://www.olostep.com/images/close-x-svgrepo-com-3.svg)\n\nOlostep Manifesto \\| [read more](https://www.olostep.com/blog/about-olostep) \u2192\n\n![](https://www.olostep.com/images/close-x-svgrepo-com-3.svg)\n\n[![](https://www.olostep.com/images/olostep-logo-cropped.svg)](https://www.olostep.com/)\n\n# Web Data Infrastructure for AI\n\nBuilt to power the Web's second user, Olostep is the best web agentic search, scraping and crawling API for AI\n\n[Start for free](https://www.olostep.com/auth) [Contact Sales](https://www.olostep.com/contact-sales)\n\nThank you! Your submission has been received!\n\nOops! Something went wrong while submitting the form.\n\n[Scrape](https://www.olostep.com/#w-tabs-0-data-w-pane-0) [Crawl](https://www.olostep.com/#w-tabs-0-data-w-pane-1) [Map](https://www.olostep.com/#w-tabs-0-data-w-pane-2) [Search](https://www.olostep.com/#w-tabs-0-data-w-pane-3) [Answer](https://www.olostep.com/#w-tabs-0-data-w-pane-4)\n\n![](https://www.olostep.com/images/svgexport-11-1.svg)\n\n## Trusted by the best startups **startups** in the world\n\n![](https://www.olostep.com/images/455e150089b14aedb083b23c8e8f157f__1_-removebg-preview.png)![](https://www.olostep.com/images/airops.png)![](https://www.olostep.com/images/podqi-logo.png)![](https://www.olostep.com/images/khoj_original-removebg-preview.png)![](https://www.olostep.com/images/svgexport-1-1.svg)![](https://www.olostep.com/images/finny_ai-removebg-preview.png)![](https://www.olostep.com/images/Logo_Contents_2025_Blue-scaled.png)![](https://www.olostep.com/images/athenahq-logo-black.png)![](https://www.olostep.com/images/CivilGrid_Logo-removebg-preview.png)![](https://www.olostep.com/images/logo.svg)![](https://www.olostep.com/images/plots_black.png)![](https://www.olostep.com/images/use-bear.avif)![](https://www.olostep.com/images/Searchable.svg)![](https://www.olostep.com/images/uman-logo.svg)![](https://www.olostep.com/images/verisave-logo.png)![](https://www.olostep.com/images/relay-app-image-removebg-preview.png)![](https://www.olostep.com/images/openmart_originak-removebg-preview.png)![](https://www.olostep.com/images/profound_logo-removebg-preview.png)![](https://www.olostep.com/images/centralize-logo.png)\n\n![](https://www.olostep.com/images/455e150089b14aedb083b23c8e8f157f__1_-removebg-preview.png)![](https://www.olostep.com/images/airops.png)![](https://www.olostep.com/images/podqi-logo.png)![](https://www.olostep.com/images/khoj_original-removebg-preview.png)![](https://www.olostep.com/images/svgexport-1-1.svg)![](https://www.olostep.com/images/finny_ai-removebg-preview.png)![](https://www.olostep.com/images/Logo_Contents_2025_Blue-scaled.png)![](https://www.olostep.com/images/athenahq-logo-black.png)![](https://www.olostep.com/images/CivilGrid_Logo-removebg-preview.png)![](https://www.olostep.com/images/logo.svg)![](https://www.olostep.com/images/plots_black.png)![](https://www.olostep.com/images/use-bear.avif)![](https://www.olostep.com/images/Searchable.svg)![](https://www.olostep.com/images/uman-logo.svg)![](https://www.olostep.com/images/verisave-logo.png)![](https://www.olostep.com/images/relay-app-image-removebg-preview.png)![](https://www.olostep.com/images/openmart_originak-removebg-preview.png)![](https://www.olostep.com/images/profound_logo-removebg-preview.png)![](https://www.olostep.com/images/centralize-logo.png)\n\n## One **API** to Automate Web Data\n\nSearch, scrape, structure and monitor the whole web with one\n\nAPI. Reliable, cost-effective, scalable. Handling Billions of requests\n\n[**Monitor** \\\\\n\\\\\nSet up monitors for events happening across the Web \\\\\n\\\\\nMonitor when AirOps publishes a new blog\\\\\n\\\\\nNew blog: The Next Chapter for AirOps...](https://www.olostep.com/dashboard/research-agents)\n\n[Latest updates SpaceX...\\\\\n\\\\\nUpdates - SpaceX...\\\\\n\\\\\n**Search** \\\\\n\\\\\nSearch the web with natural language and get ranked links](https://www.olostep.com/dashboard/playground) [Roosevelt's quote of critics\\\\\n\\\\\n\"It's not the critic who \\\\\n\\\\\ncounts; not the man\"\\\\\n\\\\\n**Answer** \\\\\n\\\\\nSearch the web and get AI-powered answers](https://www.olostep.com/dashboard/playground)\n\n[Olostep - The way we're gong to ratchet up our species\\\\\n\\\\\n**Scrape** \\\\\n\\\\\nGet real-time data from websites. Clean Markdown, HTML, Screenshots, JSON...](https://www.olostep.com/dashboard/playground) [Stripe Docs\\\\\n\\\\\n**Crawl** \\\\\n\\\\\nRetrieve all pages on a site and get their contents](https://www.olostep.com/dashboard/playground) [**Batch** \\\\\n\\\\\nProcess up to 100k URLs in 5-7 minutes](https://www.olostep.com/dashboard/playground)\n\n## Built for Developers\n\nObject-oriented API, native Python and NodeJS SDK clients,\n\nmetadata support, webhook events, easy to try and easy to scale\n\n[![](https://www.olostep.com/images/search-window-2.svg)\\\\\n\\\\\n/scrapes](https://www.olostep.com/#w-tabs-1-data-w-pane-0) [![](https://www.olostep.com/images/internet.svg)\\\\\n\\\\\n/crawls](https://www.olostep.com/#w-tabs-1-data-w-pane-1) [![](https://www.olostep.com/images/map.svg)\\\\\n\\\\\n/maps](https://www.olostep.com/#w-tabs-1-data-w-pane-2) [![](https://www.olostep.com/images/apple-shortcuts-1.svg)\\\\\n\\\\\n/batches](https://www.olostep.com/#w-tabs-1-data-w-pane-3) [![](https://www.olostep.com/images/search-1.svg)\\\\\n\\\\\n/searches](https://www.olostep.com/#w-tabs-1-data-w-pane-4) [![](https://www.olostep.com/images/bubble-search.svg)\\\\\n\\\\\n/answers](https://www.olostep.com/#w-tabs-1-data-w-pane-5) [![](https://www.olostep.com/images/bell-notification.svg)\\\\\n\\\\\n/monitors](https://www.olostep.com/#w-tabs-1-data-w-pane-6)\n\nGet clean data from any URL\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-2-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-2-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-2-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6result = client.scrapes.create(\n7    url_to_scrape=\"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n8    formats=[\"markdown\", \"html\"],\n9)\n10\n11print(result.markdown_content)\n12print(result.html_content)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const result = await client.scrapes.create({\n7  url: 'https://en.wikipedia.org/wiki/Alexander_the_Great',\n8  formats: ['markdown', 'html'],\n9})\n10\n11console.log(result.markdown_content)\n12console.log(result.html_content)\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/scrapes\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"url_to_scrape\": \"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n6    \"formats\": [\"markdown\", \"html\"]\n7  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/scrapes)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nCrawl all the subpages\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-3-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNodeJS](https://www.olostep.com/#w-tabs-3-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-3-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6crawl = client.crawls.create(\n7    start_url=\"https://olostep.com\",\n8    max_pages=100,\n9    include_urls=[\"/**\"],\n10    exclude_urls=[\"/collections/**\"],\n11    include_external=False,\n12)\n13\n14print(crawl.id, crawl.status)\n15\n16# Wait for completion and iterate pages\n17for page in crawl.pages():\n18    print(page.url)\n19    content = page.retrieve([\"markdown\"])\n20    print(content.markdown_content[:200])\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const crawl = await client.crawls.create({\n7  url: 'https://olostep.com',\n8  maxPages: 100,\n9  includeUrls: ['/**'],\n10  excludeUrls: ['/collections/**'],\n11  includeExternal: false,\n12})\n13\n14console.log(crawl.id, crawl.status)\n15\n16// Wait for completion and iterate pages\n17for await (const page of crawl.pages()) {\n18  console.log(page.url)\n19  const content = await client.retrieve({ retrieveId: page.retrieve_id, formats: ['markdown'] })\n20  console.log(content.markdown_content.slice(0, 200))\n21}\n```\n\n```bash\n1# Start crawl\n2curl -s -X POST \"https://api.olostep.com/v1/crawls\" \\\n3  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n4  -H \"Content-Type: application/json\" \\\n5  -d '{\n6    \"start_url\": \"https://olostep.com\",\n7    \"max_pages\": 100,\n8    \"include_urls\": [\"/**\"],\n9    \"exclude_urls\": [\"/collections/**\"],\n10    \"include_external\": false\n11  }'\n12\n13# Check status (replace <CRAWL_ID>)\n14curl -s \"https://api.olostep.com/v1/crawls/<CRAWL_ID>\" \\\n15  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n16\n17# Get pages (replace <CRAWL_ID>)\n18curl -s \"https://api.olostep.com/v1/crawls/<CRAWL_ID>/pages\" \\\n19  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n20\n21# Retrieve content (replace <RETRIEVE_ID>)\n22curl -s -G \"https://api.olostep.com/v1/retrieve\" \\\n23  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n24  --data-urlencode \"retrieve_id=<RETRIEVE_ID>\" \\\n25  --data-urlencode \"formats=markdown\"\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/crawls)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nGet all the URLs on a website\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-4-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-4-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-4-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6sitemap = client.maps.create(\n7    url=\"https://docs.olostep.com\",\n8    include_urls=[\"/features/**\"],\n9    top_n=100,\n10)\n11\n12print(f\"Map ID: {sitemap.id}\")\n13\n14# Iterate all URLs (handles pagination automatically)\n15for url in sitemap.urls():\n16    print(url)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const map = await client.maps.create({\n7  url: 'https://docs.olostep.com',\n8  includeUrls: ['/features/**'],\n9  topN: 100,\n10})\n11\n12console.log(`Map ID: ${map.id}`)\n13\n14// Iterate all URLs (handles pagination automatically)\n15for await (const url of map.urls()) {\n16  console.log(url)\n17}\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/maps\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"url\": \"https://docs.olostep.com\",\n6    \"include_urls\": [\"/features/**\"],\n7    \"top_n\": 100\n8  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/maps)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nProcess up to 10k URLs in one batch. Get results in 5-8 mins\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-5-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-5-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-5-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6batch = client.batches.create(\n7    urls=[\\\n8        {\"custom_id\": \"item-1\", \"url\": \"https://www.google.com/search?q=stripe&gl=us&hl=en\"},\\\n9        {\"custom_id\": \"item-2\", \"url\": \"https://www.google.com/search?q=paddle&gl=us&hl=en\"},\\\n10    ],\n11    parser=\"@olostep/google-search\",\n12)\n13\n14print(batch.id, batch.status)\n15\n16# Wait and iterate results (auto-waits for completion)\n17for item in batch.items():\n18    content = item.retrieve([\"json\"])\n19    print(item.url, item.custom_id)\n20    print(content.json_content)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const batch = await client.batches.create([\\\n7  { url: 'https://www.google.com/search?q=stripe&gl=us&hl=en', customId: 'item-1' },\\\n8  { url: 'https://www.google.com/search?q=paddle&gl=us&hl=en', customId: 'item-2' },\\\n9], {\n10  parser: '@olostep/google-search',\n11})\n12\n13console.log(batch.id, batch.total_urls)\n14\n15// Wait and iterate results (auto-waits for completion)\n16for await (const item of batch.items()) {\n17  const content = await item.retrieve(['json'])\n18  console.log(item.url, item.custom_id)\n19  console.log(content.json_content)\n20}\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/batches\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"items\": [\\\n6      {\"custom_id\": \"item-1\", \"url\": \"https://www.google.com/search?q=stripe&gl=us&hl=en\"},\\\n7      {\"custom_id\": \"item-2\", \"url\": \"https://www.google.com/search?q=paddle&gl=us&hl=en\"}\\\n8    ],\n9    \"parser\": {\"id\": \"@olostep/google-search\"}\n10  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/batches)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nSemantically search the Web\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-6-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-6-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-6-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6search = client.searches.create(\"Latest updates with SpaceX\")\n7\n8print(search.id, len(search.links))\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const search = await client.searches.create('Latest updates with SpaceX')\n7\n8console.log(search.id, search.links.length)\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/searches\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"query\": \"Latest updates with SpaceX\"\n6  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/search)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nGet answers from the Web\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-7-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-7-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-7-data-w-pane-2)\n\n```python\n1# pip install olostep\n2from olostep import Olostep\n3\n4client = Olostep(api_key=\"YOUR_REAL_KEY\")\n5\n6answer = client.answers.create(\n7    task=\"What does Olostep do and what is its core offering?\",\n8    json_format={\"company\": \"\", \"what_it_does\": \"\", \"core_offering\": \"\"},\n9)\n10\n11print(answer.json_content)\n12print(answer.sources)\n```\n\n```javascript\n1// npm i olostep\n2import Olostep from 'olostep'\n3\n4const client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n5\n6const answer = await client.answers.create({\n7  task: 'What does Olostep do and what is its core offering?',\n8  jsonFormat: { company: '', what_it_does: '', core_offering: '' },\n9})\n10\n11console.log(answer.json_content)\n12console.log(answer.sources)\n```\n\n```bash\n1curl -s -X POST \"https://api.olostep.com/v1/answers\" \\\n2  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n3  -H \"Content-Type: application/json\" \\\n4  -d '{\n5    \"task\": \"What does Olostep do and what is its core offering?\",\n6    \"json\": {\"company\": \"\", \"what_it_does\": \"\", \"core_offering\": \"\"}\n7  }'\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/answers)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\nMonitor pages on a schedule and get change alerts\n\n[![](https://www.olostep.com/images/download-2.svg)\\\\\nPython](https://www.olostep.com/#w-tabs-8-data-w-pane-0) [![](https://www.olostep.com/images/download-3.svg)\\\\\nNode JS](https://www.olostep.com/#w-tabs-8-data-w-pane-1) [![](https://www.olostep.com/images/code-brackets_1.svg)\\\\\ncURL](https://www.olostep.com/#w-tabs-8-data-w-pane-2)\n\n```python\n1import requests\n2import json\n3\n4API_KEY = \"<YOUR_API_KEY>\"\n5API_URL = \"https://api.olostep.com/v1\"\n6\n7# Create a monitor\n8payload = {\n9    \"query\": \"Alert me when Tesla stock price is above $500\",\n10    \"frequency\": \"every hour\",\n11    \"email\": \"alerts@example.com\"\n12}\n13\n14headers = {\n15    \"Authorization\": f\"Bearer {API_KEY}\",\n16    \"Content-Type\": \"application/json\"\n17}\n18\n19response = requests.post(f\"{API_URL}/monitors\", headers=headers, json=payload)\n20monitor = response.json()\n21monitor_id = monitor['id']\n22\n23print(f\"Monitor created: {monitor_id}\")\n24print(f\"Status: {monitor['status']}\")\n25\n26# List all monitors\n27monitors = requests.get(f\"{API_URL}/monitors\", headers=headers).json()\n28for m in monitors['monitors']:\n29    print(f\"{m['id']}: {m['url']} ({m['frequency']})\")\n30\n31# Get monitor details\n32details = requests.get(f\"{API_URL}/monitors/{monitor_id}\", headers=headers).json()\n33print(json.dumps(details, indent=2))\n34\n35# Delete a monitor\n36requests.delete(f\"{API_URL}/monitors/{monitor_id}\", headers=headers)\n37print(f\"Monitor {monitor_id} deleted\")\n```\n\n```javascript\n1const API_URL = 'https://api.olostep.com/v1'\n2const headers = {\n3  'Authorization': 'Bearer <YOUR_API_KEY>',\n4  'Content-Type': 'application/json'\n5}\n6\n7// Create a monitor\n8const res = await fetch(`${API_URL}/monitors`, {\n9  method: 'POST',\n10  headers,\n11  body: JSON.stringify({\n12    query: 'Alert me when Tesla stock price is above $500',\n13    frequency: 'every hour',\n14    email: 'alerts@example.com'\n15  })\n16})\n17\n18const monitor = await res.json()\n19console.log(`Monitor created: ${monitor.id}`)\n20console.log(`Status: ${monitor.status}`)\n21\n22// List all monitors\n23const monitors = await fetch(`${API_URL}/monitors`, { headers }).then(r => r.json())\n24monitors.monitors.forEach(m => console.log(`${m.id}: ${m.url} (${m.frequency})`))\n25\n26// Get monitor details\n27const details = await fetch(`${API_URL}/monitors/${monitor.id}`, { headers }).then(r => r.json())\n28console.log(details)\n29\n30// Delete a monitor\n31await fetch(`${API_URL}/monitors/${monitor.id}`, { method: 'DELETE', headers })\n32console.log(`Monitor ${monitor.id} deleted`)\n```\n\n```bash\n1# Create a monitor\n2curl -s -X POST \"https://api.olostep.com/v1/monitors\" \\\n3  -H \"Authorization: Bearer <YOUR_API_KEY>\" \\\n4  -H \"Content-Type: application/json\" \\\n5  -d '{\n6    \"query\": \"Track changes in product pricing and stock information\",\n7    \"url\": \"https://example.com/products/widget-pro\",\n8    \"frequency\": \"daily\",\n9    \"email\": \"alerts@example.com\"\n10  }'\n11\n12# List all monitors\n13curl -s \"https://api.olostep.com/v1/monitors\" \\\n14  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n15\n16# Get monitor details (replace <MONITOR_ID>)\n17curl -s \"https://api.olostep.com/v1/monitors/<MONITOR_ID>\" \\\n18  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n19\n20# Delete a monitor (replace <MONITOR_ID>)\n21curl -s -X DELETE \"https://api.olostep.com/v1/monitors/<MONITOR_ID>\" \\\n22  -H \"Authorization: Bearer <YOUR_API_KEY>\"\n```\n\n[Docs\\\\\n![](https://www.olostep.com/images/arrow-up-right-1.svg)](https://docs.olostep.com/features/monitors)\n\n[![](https://www.olostep.com/images/copy-1.svg)![](https://www.olostep.com/images/check.svg)](https://www.olostep.com/#)\n\n## /scrapes\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nTurn any URL into LLM-ready Markdown, HTML, screenshots, PDFs, or structured JSON. Handle JS-rendered pages, actions, and extraction workflows without maintaining browsers, proxies, or brittle scrapers.\n\n## /crawls\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nCrawl websites at scale, collect content from subpages, control depth and URL patterns, and retrieve clean HTML or Markdown for indexing, enrichment, RAG, and AI workflows.\n\n## /maps\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nDiscover every URL on a website using sitemaps and on-page links. Filter by path patterns, paginate large results, and prepare clean URL lists for SEO, crawls, and batches.\n\n## /batches\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nProcess up to 10k concurrent URLs in a single batch in 5-8 mins to get clean web data and aggregate content. Run many batches in parallel to scale to millions of concurrent requests.\n\n## /searches\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nAsk natural-language questions and get AI-generated answers grounded in web sources. Return validated data in the JSON shape you want, with NOT\\_FOUND when facts cannot be verified.\n\n## /answers\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nCreate scheduled web monitors from natural-language instructions. Track for changes on a single page or across the web, deltas, extract structured insights, and get chance notifications by email, webhook, or SMS.\n\n## /monitors\n\n![](https://www.olostep.com/images/nav-arrow-down-1.svg)\n\nTurn recurring website extraction into fast, cost-efficient structured JSON. Use pre-built parsers or create custom parsers for deterministic data pipelines. Use in conjunction with scrapes, crawls, and batches.\n\n[/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-0) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-1) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-2) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-3) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-4) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-5) [/Scrapes\\\\\n\\\\\n![](https://www.olostep.com/images/nav-arrow-down.svg)\\\\\n\\\\\nLorem ipsum dolor sit amet, consectetur adipiscing elit. Donec velit dolor, malesuada non leo ut, mattis maximus sem](https://www.olostep.com/#w-tabs-9-data-w-pane-6)\n\n[**Flexible** \\\\\n\\\\\nGet data as HTML, Markdown, text, PDF, JSON, screenshots, or raw bytes.\\\\\n\\\\\n**Structured** \\\\\n\\\\\nExtract clean, structured data using parsers or AI-powered LLM extraction.\\\\\n\\\\\n**Managed** \\\\\n\\\\\nStop managing headless browsers, proxies, and CAPTCHAs. \\\\\n\\\\\nYour Request\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nhttps:/olostep.com/\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nExtract the legal name, mission, and features\\\\\n\\\\\nScraping Infrastructure\\\\\n\\\\\n![](https://www.olostep.com/images/window-check.svg)\\\\\n\\\\\nBrowsers\\\\\n\\\\\n![](https://www.olostep.com/images/reload-window.svg)\\\\\n\\\\\nProxies\\\\\n\\\\\n![](https://www.olostep.com/images/puzzle.svg)\\\\\n\\\\\nCAPTCHAs\\\\\n\\\\\nStructured Result\\\\\n\\\\\n![](https://www.olostep.com/images/page.svg)\\\\\n\\\\\nFormats\\\\\n\\\\\nHTML\\\\\n\\\\\nMarkdown\\\\\n\\\\\nText\\\\\n\\\\\nPDF\\\\\n\\\\\nRaw Bytes\\\\\n\\\\\nScreenshot\\\\\n\\\\\n![](https://www.olostep.com/images/code-brackets-1.svg)\\\\\n\\\\\nStructured Data\\\\\n\\\\\n{ \\\\\n\\\\\n \u00a0\u00a0\u00a0\"name\": 'Olostep Technologies', \\\\\n\\\\\n\u00a0\u00a0\u00a0\u00a0\"mission\": 'Build infrastruct...', \\\\\n\\\\\n\u00a0\u00a0\u00a0\u00a0\"features\": 'monitors, batch...'\\\\\n\\\\\n}](https://www.olostep.com/playground)\n\n[**Scalable** \\\\\n\\\\\nBuilt for large-scale crawling across blogs, documentation, and websites.\\\\\n\\\\\n**Controlled** \\\\\n\\\\\nDefine crawl depth and target only URLs matching specific patterns.\\\\\n\\\\\n**Notified** \\\\\n\\\\\nReceive webhook alerts automatically when crawling jobs complete.\\\\\n\\\\\nhttps://docs.olostep.com/\\\\\n\\\\\nDepth 1\\\\\n\\\\\nDepth 2\\\\\n\\\\\nDepth 3\\\\\n\\\\\nCrawled/ Included\\\\\n\\\\\nExcluded by Rules](https://www.olostep.com/playground?q=crawl)\n\n[**Complete** \\\\\n\\\\\nDiscover all URLs across a website, including pages beyond sitemaps.\\\\\n\\\\\n**Customizable** \\\\\n\\\\\nInclude or exclude paths using flexible URL pattern matching.\\\\\n\\\\\n**Insightful** \\\\\n\\\\\nExplore website structure and identify URLs worth scraping next.\\\\\n\\\\\nhttps://www.olostep.com/\\\\\n\\\\\n![](https://www.olostep.com/images/folder-1.svg)\\\\\n\\\\\n/docs\\\\\n\\\\\n/get-started/welcome\\\\\n\\\\\n/features/maps\\\\\n\\\\\n/integrations/n8n\\\\\n\\\\\n....\\\\\n\\\\\n![](https://www.olostep.com/images/post.svg)\\\\\n\\\\\n/blog\\\\\n\\\\\n/monitors-api\\\\\n\\\\\n/parsers-vs-llm\\\\\n\\\\\n/olostep-orthogonal\\\\\n\\\\\n....\\\\\n\\\\\n![](https://www.olostep.com/images/shop.svg)\\\\\n\\\\\n/store\\\\\n\\\\\n/google-search\\\\\n\\\\\n/brave-search\\\\\n\\\\\n/bing-search\\\\\n\\\\\n....\\\\\n\\\\\n![](https://www.olostep.com/images/code-brackets-square.svg)\\\\\n\\\\\n/api-ref\\\\\n\\\\\n/scrapes/create\\\\\n\\\\\n/batches/create\\\\\n\\\\\n/crawls/create\\\\\n\\\\\n....](https://www.olostep.com/playground?q=map)\n\n[**Concurrent** \\\\\n\\\\\nProcess up to 10k concurrent URLs in a single batch in 5-8 mins.\\\\\n\\\\\n**Massive** \\\\\n\\\\\nRun multiple batches simultaneously to process millions of URLs.\\\\\n\\\\\n**Unique** \\\\\n\\\\\nUnique feature on the market purpose-built for large-scale concurrent processing.\\\\\n\\\\\nInput: URLs\\\\\n\\\\\n1,000,000\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-1\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-2\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-3\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-4\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-5\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-6\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nolostep.com/pg-n\\\\\n\\\\\nConcurrent Batch Processing\\\\\n\\\\\nBatch 1\\\\\n\\\\\nProcessing\\\\\n\\\\\nBatch 2\\\\\n\\\\\nProcessing\\\\\n\\\\\nBatch 3\\\\\n\\\\\nProcessing\\\\\n\\\\\nBatch n\\\\\n\\\\\nProcessing\\\\\n\\\\\n![](https://www.olostep.com/images/check-circle.svg)\\\\\n\\\\\nURLs Processed\\\\\n\\\\\n1,000,000\\\\\n\\\\\n![](https://www.olostep.com/images/apple-shortcuts-2.svg)\\\\\n\\\\\nConcurrent Batches\\\\\n\\\\\n128\\\\\n\\\\\n![](https://www.olostep.com/images/dashboard-speed.svg)\\\\\n\\\\\nThroughput\\\\\n\\\\\n48,752 /min\\\\\n\\\\\n![](https://www.olostep.com/images/clock.svg)\\\\\n\\\\\nTotal time\\\\\n\\\\\n20m 14s](https://www.olostep.com/playground?q=batch)\n\n[**Natural** \\\\\n\\\\\nGet relevant links with titles and descriptions using natural language queries.\\\\\n\\\\\n**Expandable** \\\\\n\\\\\nCombine search and scraping into a single API call for faster data retrieval.\\\\\n\\\\\n**Focused** \\\\\n\\\\\nInclude or exclude domains to refine search results precisely.\\\\\n\\\\\nYour Query\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nBest AI Agents Framework 2026\\\\\n\\\\\nInclude\\\\\n\\\\\nLangGraph\\\\\n\\\\\nCrewAi\\\\\n\\\\\nExclude\\\\\n\\\\\nOpenAI SDK\\\\\n\\\\\nAutoGen\\\\\n\\\\\n![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nLangGraph\\\\\n\\\\\n![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nAutoGen\\\\\n\\\\\n![](https://www.olostep.com/images/170677839.png)\\\\\n\\\\\nCrewAI\\\\\n\\\\\n![](https://www.olostep.com/images/2.jpg)\\\\\n\\\\\nOpenAI SDK\\\\\n\\\\\nStructured Results\\\\\n\\\\\n![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nhttps:/langchain...\\\\\n\\\\\nLangGraph\\\\\n\\\\\nBalance agent control with agency...\\\\\n\\\\\n![](https://www.olostep.com/images/170677839.png)\\\\\n\\\\\nhttps://crewai.com/\\\\\n\\\\\nCrewAI\\\\\n\\\\\nThe open platform that accelerates agent...](https://www.olostep.com/playground?q=search)\n\n[**Grounded** \\\\\n\\\\\nSearch the web and generate answers from real-world sources.\\\\\n\\\\\n**Structured** \\\\\n\\\\\nReturn validated results and sources in the exact JSON shape you need.\\\\\n\\\\\n**Enrich** \\\\\n\\\\\nEnhance products, datasets, and spreadsheets with web data.\\\\\n\\\\\nYour Question\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nFind YC startups building voice AI\\\\\n\\\\\nProcessing Layer\\\\\n\\\\\n![](https://www.olostep.com/images/search_1.svg)\\\\\n\\\\\nSearch\\\\\n\\\\\n![](https://www.olostep.com/images/spark.svg)\\\\\n\\\\\nClean\\\\\n\\\\\n![](https://www.olostep.com/images/check-circle.svg)\\\\\n\\\\\nValidate\\\\\n\\\\\nStructured Results\\\\\n\\\\\n{\"results\": \\[\\\\\n\\\\\n \u00a0 \u00a0{\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"company\": \"Retell AI\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"yc\\_batch\": \"W24\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"category\": \"Voice Agents\"\\\\\n\\\\\n \u00a0 \u00a0},\\\\\n\\\\\n \u00a0 \u00a0{\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"company\": \"Vapi\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"yc\\_batch\": \"W21\",\\\\\n\\\\\n \u00a0 \u00a0 \u00a0\"category\": \"Voice Infrastructure\"\\\\\n\\\\\n \u00a0 \u00a0}\\\\\n\\]}](https://www.olostep.com/playground?q=answer)\n\n[**Scheduled** \\\\\n\\\\\nTrack website changes automatically on a recurring schedule.\\\\\n\\\\\n**Alerts** \\\\\n\\\\\nGet notified via email, SMS, webhooks, or custom channels.\\\\\n\\\\\n**Deterministic** \\\\\n\\\\\nCreate monitors with natural language and run them deterministically.\\\\\n\\\\\nSetup to monitor\\\\\n\\\\\n![](https://www.olostep.com/images/sparks.svg)\\\\\n\\\\\nMonitor the OpenAI pricing page & alert me if any pricing changes\\\\\n\\\\\nSchedule\\\\\n\\\\\nEvery 6 hours\\\\\n\\\\\nAlert Delivered\\\\\n\\\\\nChange Summary\\\\\n\\\\\n$20 -> $22 (+$2)\\\\\n\\\\\n![](https://www.olostep.com/images/mail_1.svg)\\\\\n\\\\\nEmail\\\\\n\\\\\n![](https://www.olostep.com/images/link-1.svg)\\\\\n\\\\\nWebhook\\\\\n\\\\\n![](https://www.olostep.com/images/chat-bubble.svg)\\\\\n\\\\\nSms](https://www.olostep.com/dashboard/monitors)\n\n## Data tailored to your industry\n\nSee how Olostep powers AI platforms, sales lead enrichment, deep research, competitive intelligence, and SEO teams with one API.\n\n### Deep Search\n\nAccess custom, hyper-specialized B2B indexes for your industry to search and extract comprehensive data beyond what general web indexes cover\n\nAdd Enrichment\n\n![](https://www.olostep.com/images/group.svg)\n\nSearch for persons at company\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/mail.svg)\n\nFind business emails\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/coins.svg)\n\nFind annual revenues\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/suitcase.svg)\n\nSearch job openings\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/map-pin.svg)\n\nFind address\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Recruiting\n\nIdentify, research, and validate candidates faster with intelligence and data aggregated from top-quality profiles and specialist web sources.\n\nRecruiting Pipeline\n\n![](https://www.olostep.com/images/search-engine.svg)\n\nSourcing\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/filter.svg)\n\nPreprocessing\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/label.svg)\n\nLabeling\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/stats-down-square.svg)\n\nEvaluation\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/rocket.svg)\n\nDeployment\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Power AI applications\n\nGet clean, structured data from any website as markdown, html, screenshot, etc. to power your AI application and workflows\n\nExtract Content\n\n![](https://www.olostep.com/images/html5.svg)\n\nExtract as Markdown / HTML\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/code-brackets.svg)\n\nExtract as JSON\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/code.svg)\n\nExtract Code\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/multiple-pages.svg)\n\nExtract as Text / PDF\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/screenshot.svg)\n\nExtract as Screenshot\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Monitor the Web\n\nMonitor any webpage for DOM changes, stock availability, price changes, job openings or fresh content. Run automatically on a schedule and get alerted\n\nMonitor Webpage\n\n![](https://www.olostep.com/images/candlestick-chart.svg)\n\nMonitor pricing & stock\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/eye.svg)\n\nMonitor content & DOM changes\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/suitcase.svg)\n\nMonitor job openings\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/leaderboard-star.svg)\n\nMonitor reviews & ratings\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/alarm.svg)\n\nSet schedule & alerts\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n### Automate data pipelines\n\nAutomate complex data pipelines with the /agents endpoint through natural language prompts. You can also pass your own internal knowledge as context\n\nNatural language to data pipelines\n\n![](https://www.olostep.com/images/search-engine.svg)\n\nPower your search product\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/eye.svg)\n\nTrack portfolio companies\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/coins.svg)\n\nBuild market intelligence\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/stats-down-square.svg)\n\nResearch companies at scale\n\n![](https://www.olostep.com/images/Group-2.svg)\n\n![](https://www.olostep.com/images/suitcase.svg)\n\nAutomate GTM and recruiting\n\n![](https://www.olostep.com/images/Group-2.svg)\n\nAutomate data pipelines\n\nAutomate complex data pipelines with the /agents endpoint through natural language prompts. You can also pass your own internal knowledge as context\n\n![](https://www.olostep.com/images/image-5.png)\n\n[![](https://www.olostep.com/images/nav-arrow-left.svg)](https://www.olostep.com/#)[![](https://www.olostep.com/images/nav-arrow-right.svg)](https://www.olostep.com/#)\n\n## Pricing that Makes Sense\n\nMost cost-effective web data API on the market\n\n**No credit card required**\n\n### Trial\n\n$0\n\nIncludes:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n500 successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\nAll requests are JS rendered + utilizing residential IP addresses\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\nLow rate limits\n\n[Get started](https://www.olostep.com/dashboard/plans)\n\nCOST/1K$1.800\n\n### Starter\n\n$9\n\n/ month\n\nEverything in Free, Plus:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n5000 successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n150 concurrent requests\n\n[Purchase now](https://www.olostep.com/dashboard/plans)\n\nCOST/1K**$0.495**\n\n### Standard\n\n$99\n\n/ month\n\nEverything in Starter, Plus:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n200K successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n500 concurrent requests\n\n[Purchase now](https://www.olostep.com/dashboard/plans)\n\nCOST/1K$ **0.399**\n\n### Scale\n\n$399\n\n/ month\n\nEverything in Standard, Plus:\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\n1 Million successful requests\n\n![](https://www.olostep.com/images/check-svgrepo-com.svg)\n\nAI-powered Browser Automations\n\n[Purchase now](https://www.olostep.com/dashboard/plans)\n\n## Top-ups\n\nHave spiky usage or don't like subscriptions?\n\nYou can buy credit packs. They are valid for 6 months.\n\nCredit pack\n\n### 10k credits\n\n**$20**\n\n[Purchase Credit Pack](https://www.olostep.com/dashboard/plans)\n\nCredit pack\n\n### 250k credits\n\n**$200**\n\n[Purchase Credit Pack](https://www.olostep.com/dashboard/plans)\n\nCredit pack\n\n### 2M credits\n\n**$1000**\n\n[Purchase Credit Pack](https://www.olostep.com/dashboard/plans)\n\n### Enterprise\n\nHundreds of millions of credits with enterprise-grade reliability. We offer custom discounts\n\n[Contact Sales](https://www.olostep.com/contact-sales)\n\n## Trusted by Amazing Teams Building the Future of AI\n\n![](https://www.olostep.com/images/michelle-pic.jpg)\n\nMichelle Julia\n\nCo-founder & CEO Aurium\n\nOlostep is the best!!! We automated entire data pipelines with just a prompt\n\n![](https://www.olostep.com/images/richard-he.avif)\n\nRichard He\n\nCo-founder & CEO Openmart\n\nOlostep has become the default Web Layer infrastructure for our company\n\n![](https://www.olostep.com/images/mx-pic.jpeg)\n\nMax Brodeur-Urbas\n\nCo-founder & CEO Gumloop\n\nOlostep works like a charm! And your customer service is exceptional\n\n![](https://www.olostep.com/images/rob-pic.jpeg)\n\nRob Hayes\n\nCo-founder Merchkit\n\nOlostep lets us turn any website into an API. Great product, great people\n\n![](https://www.olostep.com/images/brandon-civilgrid.webp)\n\nBrandon Cohen\n\nCo-founder & CTO CivilGrid\n\nI highly recommend Olostep, great product!\n\n![](https://www.olostep.com/images/michelle-pic.jpg)\n\nMichelle Julia\n\nCo-founder & CEO Aurium\n\nOlostep is the best!!! We automated entire data pipelines with just a prompt\n\n![](https://www.olostep.com/images/richard-he.avif)\n\nRichard He\n\nCo-founder & CEO Openmart\n\nOlostep has become the default Web Layer infrastructure for our company\n\n![](https://www.olostep.com/images/mx-pic.jpeg)\n\nMax Brodeur-Urbas\n\nCo-founder & CEO Gumloop\n\nOlostep works like a charm! And your customer service is exceptional\n\n![](https://www.olostep.com/images/rob-pic.jpeg)\n\nRob Hayes\n\nCo-founder Merchkit\n\nOlostep lets us turn any website into an API. Great product, great people\n\n![](https://www.olostep.com/images/brandon-civilgrid.webp)\n\nBrandon Cohen\n\nCo-founder & CTO CivilGrid\n\nI highly recommend Olostep, great product!\n\n![](https://www.olostep.com/images/berk-pic.jpeg)\n\n[Berk Serbetcioglu](https://www.berkserbetcioglu.com/)\n\nCo-founder & CEO Gedd.it\n\nWe verify coupon codes at scale. Love Olostep. It works on any e-commerce\n\n![](https://www.olostep.com/images/trevor-west.jpeg)\n\nTrevor West\n\nCo-founder & CEO Podqi\n\nOlostep is the best API to search, extract, and structure data from the Web. Happy to be customers\n\n![](https://www.olostep.com/images/rida_pic.jpg)\n\nRida Naveed\n\nCo-founder Zecento\n\nWe use /batches combined with parsers and it's magical how we can get structured data at large scale\n\n![](https://www.olostep.com/images/kieran-plots.jpeg)\n\nKieran V.\n\nGrowth PlotsEvents\n\nOlostep allowed us to search and structure events data across the Web\n\n![](https://www.olostep.com/images/paul-mit.jpg)\n\nPaul Mit\n\nFounder Foundbase\n\nReliable and cost-effective API for working with data. Congrats on the cool product\n\n![](https://www.olostep.com/images/berk-pic.jpeg)\n\n[Berk Serbetcioglu](https://www.berkserbetcioglu.com/)\n\nCo-founder & CEO Gedd.it\n\nWe verify coupon codes at scale. Love Olostep. It works on any e-commerce\n\n![](https://www.olostep.com/images/trevor-west.jpeg)\n\nTrevor West\n\nCo-founder & CEO Podqi\n\nOlostep is the best API to search, extract, and structure data from the Web. Happy to be customers\n\n![](https://www.olostep.com/images/rida_pic.jpg)\n\nRida Naveed\n\nCo-founder Zecento\n\nWe use /batches combined with parsers and it's magical how we can get structured data at large scale\n\n![](https://www.olostep.com/images/kieran-plots.jpeg)\n\nKieran V.\n\nGrowth PlotsEvents\n\nOlostep allowed us to search and structure events data across the Web\n\n![](https://www.olostep.com/images/paul-mit.jpg)\n\nPaul Mit\n\nFounder Foundbase\n\nReliable and cost-effective API for working with data. Congrats on the cool product\n\n## Connect Olostep to Your AI Stack\n\nOfficial Olostep integrations. Add web scraping,\n\ncrawling and AI-powered search to any tool in your stack.\n\n[![](https://www.olostep.com/images/cursor.webp)\\\\\n\\\\\nCursor](https://www.olostep.com/#) [![](https://www.olostep.com/images/claude.jpeg)\\\\\n\\\\\nClaude](https://www.olostep.com/#) [![](https://www.olostep.com/images/n8n.webp)\\\\\n\\\\\nn8n](https://www.olostep.com/#) [![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nLangChain](https://www.olostep.com/#) [![](https://www.olostep.com/images/4.jpg)\\\\\n\\\\\nApify](https://www.olostep.com/#) [![](https://www.olostep.com/images/149120496.png)\\\\\n\\\\\nMastra](https://www.olostep.com/#) [![](https://www.olostep.com/images/zapier.webp)\\\\\n\\\\\nZapier](https://www.olostep.com/#) [![](https://www.olostep.com/images/relay.webp)\\\\\n\\\\\nRelay](https://www.olostep.com/#) [![](https://www.olostep.com/images/windsurf.webp)\\\\\n\\\\\nWindsurf](https://www.olostep.com/#) [![](https://www.olostep.com/images/3.jpg)\\\\\n\\\\\nCline](https://www.olostep.com/#) [![](https://www.olostep.com/images/vscode.webp)\\\\\n\\\\\nVS Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/1.jpg)\\\\\n\\\\\nGemini CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/2.jpg)\\\\\n\\\\\nOpenAI Codex](https://www.olostep.com/#) [![](https://www.olostep.com/images/cursor.webp)\\\\\n\\\\\nCursor](https://www.olostep.com/#) [![](https://www.olostep.com/images/claude.jpeg)\\\\\n\\\\\nClaude](https://www.olostep.com/#) [![](https://www.olostep.com/images/n8n.webp)\\\\\n\\\\\nn8n](https://www.olostep.com/#) [![](https://www.olostep.com/images/langchain.webp)\\\\\n\\\\\nLangChain](https://www.olostep.com/#) [![](https://www.olostep.com/images/4.jpg)\\\\\n\\\\\nApify](https://www.olostep.com/#) [![](https://www.olostep.com/images/149120496.png)\\\\\n\\\\\nMastra](https://www.olostep.com/#) [![](https://www.olostep.com/images/zapier.webp)\\\\\n\\\\\nZapier](https://www.olostep.com/#) [![](https://www.olostep.com/images/relay.webp)\\\\\n\\\\\nRelay](https://www.olostep.com/#) [![](https://www.olostep.com/images/windsurf.webp)\\\\\n\\\\\nWindsurf](https://www.olostep.com/#) [![](https://www.olostep.com/images/3.jpg)\\\\\n\\\\\nCline](https://www.olostep.com/#) [![](https://www.olostep.com/images/vscode.webp)\\\\\n\\\\\nVS Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/1.jpg)\\\\\n\\\\\nGemini CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/2.jpg)\\\\\n\\\\\nOpenAI Codex](https://www.olostep.com/#)\n\n[![](https://www.olostep.com/images/qwen.jpg)\\\\\n\\\\\nQwen](https://www.olostep.com/#) [![](https://www.olostep.com/images/smithery.jpg)\\\\\n\\\\\nSmithery](https://www.olostep.com/#) [![](https://www.olostep.com/images/amazonaws.webp)\\\\\n\\\\\nAmazon Q](https://www.olostep.com/#) [![](https://www.olostep.com/images/ampcode.webp)\\\\\n\\\\\nAmp](https://www.olostep.com/#) [![](https://www.olostep.com/images/augment.jpg)\\\\\n\\\\\nAugment](https://www.olostep.com/#) [![](https://www.olostep.com/images/boltai.webp)\\\\\n\\\\\nBoltAI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Bun.jpg)\\\\\n\\\\\nBun](https://www.olostep.com/#) [![](https://www.olostep.com/images/deno-1.webp)\\\\\n\\\\\nDeno](https://www.olostep.com/#) [![](https://www.olostep.com/images/copilot.webp)\\\\\n\\\\\nCopilot](https://www.olostep.com/#) [![](https://www.olostep.com/images/124303983.png)\\\\\n\\\\\nCrush](https://www.olostep.com/#) [![](https://www.olostep.com/images/docker.webp)\\\\\n\\\\\nDocker](https://www.olostep.com/#) [![](https://www.olostep.com/images/jetbrains.webp)\\\\\n\\\\\nJetBrains](https://www.olostep.com/#) [![](https://www.olostep.com/images/kiro.webp)\\\\\n\\\\\nKiro](https://www.olostep.com/#) [![](https://www.olostep.com/images/qwen.jpg)\\\\\n\\\\\nQwen](https://www.olostep.com/#) [![](https://www.olostep.com/images/smithery.jpg)\\\\\n\\\\\nSmithery](https://www.olostep.com/#) [![](https://www.olostep.com/images/amazonaws.webp)\\\\\n\\\\\nAmazon Q](https://www.olostep.com/#) [![](https://www.olostep.com/images/ampcode.webp)\\\\\n\\\\\nAmp](https://www.olostep.com/#) [![](https://www.olostep.com/images/augment.jpg)\\\\\n\\\\\nAugment](https://www.olostep.com/#) [![](https://www.olostep.com/images/boltai.webp)\\\\\n\\\\\nBoltAI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Bun.jpg)\\\\\n\\\\\nBun](https://www.olostep.com/#) [![](https://www.olostep.com/images/deno-1.webp)\\\\\n\\\\\nDeno](https://www.olostep.com/#) [![](https://www.olostep.com/images/copilot.webp)\\\\\n\\\\\nCopilot](https://www.olostep.com/#) [![](https://www.olostep.com/images/124303983.png)\\\\\n\\\\\nCrush](https://www.olostep.com/#) [![](https://www.olostep.com/images/docker.webp)\\\\\n\\\\\nDocker](https://www.olostep.com/#) [![](https://www.olostep.com/images/jetbrains.webp)\\\\\n\\\\\nJetBrains](https://www.olostep.com/#) [![](https://www.olostep.com/images/kiro.webp)\\\\\n\\\\\nKiro](https://www.olostep.com/#)\n\n[![](https://www.olostep.com/images/LM.jpg)\\\\\n\\\\\nLM Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/opencode.webp)\\\\\n\\\\\nOpencode](https://www.olostep.com/#) [![](https://www.olostep.com/images/perplexity.webp)\\\\\n\\\\\nPerplexity](https://www.olostep.com/#) [![](https://www.olostep.com/images/qodo.webp)\\\\\n\\\\\nQodo Gen](https://www.olostep.com/#) [![](https://www.olostep.com/images/Roo.jpg)\\\\\n\\\\\nRoo Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/atlassian.webp)\\\\\n\\\\\nRovo Dev CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Traelogo.png)\\\\\n\\\\\nTrae](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nVisual Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/warp.webp)\\\\\n\\\\\nWarp](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nWindows](https://www.olostep.com/#) [![](https://www.olostep.com/images/zed.webp)\\\\\n\\\\\nZed](https://www.olostep.com/#) [![](https://www.olostep.com/images/Zencoder.jpg)\\\\\n\\\\\nZencoder](https://www.olostep.com/#) [![](https://www.olostep.com/images/modelcontextprotocol.webp)\\\\\n\\\\\nMCP client](https://www.olostep.com/#) [![](https://www.olostep.com/images/LM.jpg)\\\\\n\\\\\nLM Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/opencode.webp)\\\\\n\\\\\nOpencode](https://www.olostep.com/#) [![](https://www.olostep.com/images/perplexity.webp)\\\\\n\\\\\nPerplexity](https://www.olostep.com/#) [![](https://www.olostep.com/images/qodo.webp)\\\\\n\\\\\nQodo Gen](https://www.olostep.com/#) [![](https://www.olostep.com/images/Roo.jpg)\\\\\n\\\\\nRoo Code](https://www.olostep.com/#) [![](https://www.olostep.com/images/atlassian.webp)\\\\\n\\\\\nRovo Dev CLI](https://www.olostep.com/#) [![](https://www.olostep.com/images/Traelogo.png)\\\\\n\\\\\nTrae](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nVisual Studio](https://www.olostep.com/#) [![](https://www.olostep.com/images/warp.webp)\\\\\n\\\\\nWarp](https://www.olostep.com/#) [![](https://www.olostep.com/images/Windows.jpg)\\\\\n\\\\\nWindows](https://www.olostep.com/#) [![](https://www.olostep.com/images/zed.webp)\\\\\n\\\\\nZed](https://www.olostep.com/#) [![](https://www.olostep.com/images/Zencoder.jpg)\\\\\n\\\\\nZencoder](https://www.olostep.com/#) [![](https://www.olostep.com/images/modelcontextprotocol.webp)\\\\\n\\\\\nMCP client](https://www.olostep.com/#)\n\n![](https://www.olostep.com/images/flash.svg)\n\n### Connect with your AI agents\n\nA deterministic, repeatable, controllable pipeline that automates any web research workflow and pipeline exactly as you described it.\n\n![](https://www.olostep.com/images/olostep-logo-cropped.svg)\n\n![](https://www.olostep.com/images/n8n.webp)\n\n![](https://www.olostep.com/images/7NLTL1CG1JNK-180x180.jpeg)\n\n![](https://www.olostep.com/images/claude.jpeg)\n\n![](https://www.olostep.com/images/windsurf.webp)\n\n![](https://www.olostep.com/images/perplexity.webp)\n\n![](https://www.olostep.com/images/vscode.webp)\n\n![](https://www.olostep.com/images/terminal.svg)\n\n### Use the Olostep CLI\n\nMap, scrape, crawl, batch process, and generate answers directly from your terminal, with clean JSON output built for scripts, CI pipelines and AI agents.\n\n`npx -y olostep-cli@latest --help`\n\n![](https://www.olostep.com/images/ev-plug.svg)\n\n### Add Olostep to your MCP client\n\nWorks with any product that implements the Model Context Protocol: register Olostep once and call web tools from chat, agents, or IDEs that support MCP.\n\n`{\n\"mcpServers\": {\n \u00a0 \u00a0\"olostep-web\": {\n \u00a0 \u00a0 \u00a0\"command\": \"``npx``\",\n \u00a0 \u00a0 \u00a0\"args\": [\"``-y``\", \"``olostep-mcp``\"],\n \u00a0 \u00a0 \u00a0\"env\": {\n \u00a0 \u00a0 \u00a0 \u00a0\"OLOSTEP_API_KEY\": \"``YOUR_API_KEY``\"\n \u00a0 \u00a0 \u00a0}\n \u00a0 \u00a0}\n}\n}`\n\n## Ready to start?\n\nGet **clean data** for your AI from any website with Olostep\n\nMost cost-effective API. Built for scale\n\n[Start for free](https://www.olostep.com/auth) [See Pricing](https://www.olostep.com/pricing)\n\n[Are you an AI Agent? Get started here](https://www.olostep.com/agent-onboarding/SKILL.md)\n\n## Frequently asked questions\n\nProduct & Capabilities\n\nUsage & Automation\n\nPricing & Plans\n\n### Does Olostep offer a free trial?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYes, Olostep includes a free plan with 500 requests to help you test the API before upgrading. Paid plans start from $9/month and include 5,000 credits per month.\n\nThis gives teams a low-risk way to evaluate Olostep's reliability, scalability and cost-effectiveness before moving to higher-volume usage.\n\n### Can I switch plans after signing up?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYes, you can switch plans at any time. Plans are pro-rated, meaning any unused value from your current plan is carried over to your new plan.\n\nThis ensures you don't pay twice for usage you've already covered, giving you flexibility as your needs grow.\n\n### Can I ask for a refund if I don't use it?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYes. If you're not satisfied with the Olostep API or it doesn't end up being useful for your use case, you can email [info@olostep.com](mailto:info@olostep.com?subject=I%27d%20like%20a%20refund) to request a refund.\n\nIf you cancel after a period of non-use, Olostep can also refund the unused portion of your plan where applicable.\n\n### How can I pay?\n\n![](https://www.olostep.com/images/6177739448baa6e16cce1db4_icon_plus.svg)\n\nYou can pay through Stripe Payment Links. To access billing, you must be logged in to your Olostep account. Go to your dashboard, navigate to Billing & Invoices, click Manage on Stripe, and add your card details. Once your card is added, you\u2019re good to go.\n\n[![](https://www.olostep.com/images/olostep-logo-cropped.svg)](https://www.olostep.com/)\n\nInfrastructure for the Web's second user\n\n[![Olostep - Turn the Web into Clean Data for AI | Product Hunt](https://api.producthunt.com/widgets/embed-image/v1/featured.svg?post_id=1199418&theme=light&t=1788106286872)](https://www.producthunt.com/products/olostep?embed=true&utm_source=badge-featured&utm_medium=badge&utm_campaign=badge-olostep-2)\n\n[![](https://www.olostep.com/images/slack-svgrepo-com.svg)](https://olostep-users.slack.com/join/shared_invite/zt-2bfddyi8h-JzfjOgavg~98DJ1om1B5Lg#/shared-invite/email)\n\n[All services are online](https://status.olostep.com/)\n\nProduct\n\n[Playground](https://www.olostep.com/playground) [Agent](https://www.olostep.com/agent) [Orbit](https://www.olostep.com/orbit) [MCP Server](https://www.olostep.com/integrations/mcp-server) [Integrations](https://www.olostep.com/integrations) [Tools](https://www.olostep.com/tools) [Pricing](https://www.olostep.com/pricing)\n\nDevelopers\n\n[API Reference](https://docs.olostep.com/api-reference/scrapes/create) [API Endpoints](https://www.olostep.com/api-endpoints) [Scrape API](https://www.olostep.com/api-endpoints/scrapes) [Crawl API](https://www.olostep.com/api-endpoints/crawls) [Map API](https://www.olostep.com/api-endpoints/maps) [Batch API](https://www.olostep.com/api-endpoints/batches) [Search API](https://www.olostep.com/api-endpoints/searches) [Answer API](https://www.olostep.com/api-endpoints/answers) [Monitor API](https://www.olostep.com/api-endpoints/monitors) [People and Company Search API](https://www.olostep.com/company-people-search-api)\n\nSolutions\n\n[Tools](https://www.olostep.com/tools) [Connect to Slack](https://www.olostep.com/slack) [Use Cases](https://www.olostep.com/use-cases) [AI Platforms](https://www.olostep.com/use-cases/power-ai-platforms) [Deep Research](https://www.olostep.com/use-cases/deep-research) [Competitive Intelligence](https://www.olostep.com/use-cases/competitive-intelligence) [Sales Lead Enrichment](https://www.olostep.com/use-cases/sales-lead-enrichment) [SEO & GEO Teams](https://www.olostep.com/seo-ai-visibility) [Brand Protection](https://www.olostep.com/brand-protection)\n\nResources\n\n[Documentation](https://docs.olostep.com/get-started/welcome) [Blog](https://www.olostep.com/blog) [Glossary](https://www.olostep.com/glossary) [Changelog](https://www.olostep.com/changelog)\n\nCompany\n\n[About](https://www.olostep.com/about) [Careers](https://www.olostep.com/careers) [Partners](https://www.olostep.com/partners) [Contact Sales](https://www.olostep.com/contact-sales) [AI\u00a0Instructions](https://www.olostep.com/ai-instructions)\n\nLegal\n\nSocials\n\n[X](https://x.com/olostep) [LinkedIn](https://www.linkedin.com/company/olostep/) [Join Slack](https://olostep-users.slack.com/join/shared_invite/zt-2bfddyi8h-JzfjOgavg~98DJ1om1B5Lg#/shared-invite/email)\n\n![](https://www.olostep.com/images/headset-help.svg)\n\nNeed a hand with anything? [Join our Slack](https://olostep-users.slack.com/join/shared_invite/zt-2bfddyi8h-JzfjOgavg~98DJ1om1B5Lg#/shared-invite/email) or send an email to [info@olostep.com](mailto:info@olostep.com?subject=Customer%20support%20request)\n\n\u00a9 2026 Olostep Technologies, Inc.\n\nMade in Italy \\| United States",
          "metadata": {
            "ogImage": "https://cdn.prod.website-files.com/6665ee948e0a3f19d93e90c7/6a8fc7e4dc1cd02fd8615c2b_Open%20Graph%20Preview%20(1).png",
            "description": "Search, scrape, crawl, structure, and monitor the web with one API. Power AI agents, research, enrichment, RAG, and real-time data workflows.",
            "viewport": "width=device-width, initial-scale=1",
            "language": "en",
            "og:image": "https://cdn.prod.website-files.com/6665ee948e0a3f19d93e90c7/6a8fc7e4dc1cd02fd8615c2b_Open%20Graph%20Preview%20(1).png",
            "robots": "index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1",
            "generator": "Webflow",
            "title": "Web Data Infrastructure for AI Agents | Olostep",
            "twitter:image": "https://cdn.prod.website-files.com/6665ee948e0a3f19d93e90c7/6a8fc7e4dc1cd02fd8615c2b_Open%20Graph%20Preview%20(1).png",
            "favicon": "https://www.olostep.com/images/favicon.png",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-93f0d3e9a1cc",
            "sourceURL": "https://www.olostep.com/",
            "url": "https://www.olostep.com/",
            "statusCode": 200,
            "contentType": "text/html",
            "proxyUsed": "basic",
            "cacheState": "hit",
            "cachedAt": "2026-09-02T09:19:31.650Z",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://docs.olostep.com/integrations/n8n",
          "title": "Olostep + n8n",
          "description": "## [\u200b](https://docs.olostep.com/integrations/n8n#example-workflow-lead-enrichment-from-google-sheets)  Example workflow: Lead enrichment from Google Sheets\n### [\u200b](https://docs.olostep.com/integrations/n8n#step-4-add-an-openai-node)  Step 4: Add an OpenAI node\n```\nYou are a sales researcher. Based on the company website content below, extract:\n1. Industry (one phrase, e.g. \"B2B SaaS\", \"E-commerce\", \"Healthcare\")\n2. One-sentence company description (max 20 words)\n3. Estimated company size (Startup / SMB / Mid-market / Enterprise)\n\nReturn only a JSON object with keys: industry, description, company_size.\n\nWebsite content:\n{{ $json.markdownContent }}\n```\n\n## [\u200b](https://docs.olostep.com/integrations/n8n#troubleshooting)  Troubleshooting\n```\n[\\\n  { \"url\": \"https://example.com/page-1\", \"custom_id\": \"p1\" },\\\n  { \"url\": \"https://example.com/page-2\", \"custom_id\": \"p2\" }\\\n]\n```",
          "position": 2,
          "markdown": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/integrations/n8n#content-area)\n\nThe verified [Olostep Web Scraper node](https://n8n.io/integrations/olostep-web-scraper/) gives you six operations inside n8n\u2019s visual builder: scrape a URL, search the web, get AI answers, batch-scrape thousands of URLs, crawl a site, or map all its links.[View on n8n \u2192](https://n8n.io/integrations/olostep-web-scraper/)\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#before-you-start)  Before you start\n\n- **An Olostep account with an API key:** [get one free](https://olostep.com/dashboard), no credit card required. Your first 500 credits are included.\n- **n8n running:** either [n8n Cloud](https://n8n.io/cloud/) or a self-hosted instance. Community nodes must be enabled (they are by default on most setups).\n- **No coding required:** everything in this guide is done through n8n\u2019s visual editor.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#setup)  Setup\n\n1\n\nSearch for the Olostep node\n\nOpen any workflow, click **+**, and search for **Olostep**. Select **Olostep Web Scraper** from the results.![Search for Olostep in the n8n node picker](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-1.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=dee8769fcd2a85836d8fb16c12f577ba)\n\n2\n\nInstall the node\n\nClick the result to open the Node details panel, then click **Install node**. n8n will install `n8n-nodes-olostep` and prompt you to restart. Do that before continuing.![Olostep Web Scraper node details with Install node button](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-2.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=2721be5e49905b7428c6a082ed73dd18)\n\nIf **Community Nodes** is disabled for your workspace, an admin needs to enable it first. See the [n8n community nodes guide](https://docs.n8n.io/integrations/community-nodes/installation/).\n\n3\n\nAdd your API key\n\nOpen the Olostep node in your workflow, click **Set up Credential** (in the Parameters tab), add your API key, and click **Save**.![Olostep credentials form in n8n with API Key field](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-3.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=ce6d9f0d70ee4f4a8f167bf1909b7c8d)Get your key from the [Olostep dashboard \u2192](https://olostep.com/dashboard)\n\n4\n\nWire it up and run\n\nConnect the Olostep node to a trigger and any downstream steps, then execute your workflow.![n8n workflow canvas with Schedule Trigger connected to Olostep node](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/step-4.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=bf0a711ca890d85ed1b8915c50035697)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#actions)  Actions\n\n## Scrape Website\n\nPull content from any URL as Markdown, HTML, JSON, or plain text. Handles JS-rendered pages with optional wait times and country targeting.\n\n## Search\n\nRun a web search and get structured results (titles, URLs, and snippets) as JSON.\n\n## Answers (AI)\n\nAsk a natural-language question and get an answer with cited sources. Useful before LLM nodes when you need grounded responses.\n\n## Batch Scrape URLs\n\nSubmit up to 10,000 URLs in one job, processed in parallel. Returns a `batch_id`; retrieve results asynchronously.\n\n## Create Crawl\n\nStart from a URL, follow links, and scrape all subpages. Good for docs sites, blogs, or full-site ingestion. Returns a `crawl_id`.\n\n## Create Map\n\nGet every URL on a site without scraping content. Use it for discovery before a batch job. Returns a `map_id`.\n\n**Batch, Crawl, and Map are async.** Store the returned ID and use a Wait node or a second workflow to retrieve results once processing completes.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#example-workflow-lead-enrichment-from-google-sheets)  Example workflow: Lead enrichment from Google Sheets\n\n**What it does:** When you paste a company URL into a Google Sheet, this workflow automatically scrapes the company\u2019s website, extracts key information with an AI node, and writes the results back to the same row, turning a blank spreadsheet into a filled-out lead database.**Nodes used:** Google Sheets trigger \u2192 Olostep Scrape Website \u2192 OpenAI \u2192 Code \u2192 Google Sheets update![Lead enrichment workflow in n8n: Google Sheets trigger connected to Olostep, OpenAI, Code, and Google Sheets update nodes](https://mintcdn.com/olostep-58/dxAYTC_H6gJN108B/images/integrations/n8n/workflow-example.png?fit=max&auto=format&n=dxAYTC_H6gJN108B&q=85&s=60e725c70903c2b3cf6918f2ca486e70)\n\n* * *\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-1-set-up-your-google-sheet)  Step 1: Set up your Google Sheet\n\nCreate a sheet with these columns: `Company URL`, `Industry`, `Description`, `Company Size`, `Enriched`. The workflow reads from `Company URL` and fills in the rest.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-2-add-a-google-sheets-trigger)  Step 2: Add a Google Sheets trigger\n\nIn n8n, add a **Google Sheets** trigger node. Set the event to **Row Added**, point it at your sheet, and set it to watch the `Company URL` column. Now every time you paste a new URL into the sheet, this workflow fires.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-3-add-olostep-scrape-website)  Step 3: Add Olostep Scrape Website\n\nConnect an **Olostep Web Scraper** node after the trigger. Set:\n\n- **Action:** Scrape Website\n- **URL:**`{{ $json[\"Company URL\"] }}` (pulls the URL from the new row)\n- **Output Format:** Markdown\n\nMarkdown works best here because it strips navigation, ads, and boilerplate. The AI node in the next step gets clean prose about the company instead of raw HTML noise.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-4-add-an-openai-node)  Step 4: Add an OpenAI node\n\nConnect an **OpenAI** node. Set the model to `gpt-4o-mini` (fast and cheap for extraction tasks) and use this prompt:\n\n```\nYou are a sales researcher. Based on the company website content below, extract:\n1. Industry (one phrase, e.g. \"B2B SaaS\", \"E-commerce\", \"Healthcare\")\n2. One-sentence company description (max 20 words)\n3. Estimated company size (Startup / SMB / Mid-market / Enterprise)\n\nReturn only a JSON object with keys: industry, description, company_size.\n\nWebsite content:\n{{ $json.markdownContent }}\n```\n\nThe `markdownContent` field is what Olostep returns from the scrape, as clean plain text.\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#step-5-parse-the-ai-response-and-write-back)  Step 5: Parse the AI response and write back\n\nAdd a **Code** node to parse the JSON from OpenAI:\n\n```\nconst parsed = JSON.parse($input.first().json.message.content);\nreturn [{ json: parsed }];\n```\n\nThen connect a **Google Sheets** node set to **Update Row**. Map the columns:\n\n- `Industry` \u2192 `{{ $json.industry }}`\n- `Description` \u2192 `{{ $json.description }}`\n- `Company Size` \u2192 `{{ $json.company_size }}`\n- `Enriched` \u2192 `Yes`\n\n### [\u200b](https://docs.olostep.com/integrations/n8n\\#what-you-get)  What you get\n\nPaste a URL like `https://notion.so` into your sheet, and within ~10 seconds the row fills in:\n\n| Company URL | Industry | Description | Company Size | Enriched |\n| --- | --- | --- | --- | --- |\n| [https://notion.so](https://notion.so/) | Productivity SaaS | All-in-one workspace for notes, docs, and databases | Mid-market | Yes |\n\nFrom here you can extend this workflow: add a Slack notification when enrichment completes, filter by industry before writing back, or replace Google Sheets with HubSpot to update contacts directly.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#templates)  Templates\n\nReady-to-import n8n workflows built with Olostep:\n\n[**Crawl docs \u2192 AI knowledge base** \\\\\n\\\\\nCrawl documentation sites with Olostep and structure the output into an AI-ready knowledge base.](https://www.n8n.io/workflows/13436-crawl-documentation-sites-and-build-an-ai-knowledge-base-with-olostep/)\n\n[**Google Maps leads \u2192 decision-maker enrichment** \\\\\n\\\\\nScrape business leads from Google Maps and enrich them with decision-maker details.](https://n8n.io/workflows/11086-scrape-business-leads-from-google-maps-and-extract-decision-maker-info-with-olostep/)\n\n[**Mine user complaints \u2192 insight report** \\\\\n\\\\\nAnalyze complaints with Olostep + Gemini and generate structured insight reports in Google Docs.](https://www.n8n.io/workflows/13435-mine-user-complaints-and-generate-insight-reports-with-olostep-gemini-and-google-docs/)\n\n[**Amazon product extraction \u2192 Google Sheets** \\\\\n\\\\\nExtract Amazon product URLs and metadata with Olostep, then sync the results to Sheets.](https://n8n.io/workflows/11158-extract-amazon-product-data-to-sheets-with-olostep-api/)\n\n[Browse all Olostep workflows on n8n.io \u2192](https://n8n.io/workflows/?q=olostep)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#parsers)  Parsers\n\nAdd a parser ID to the **Parser** field on any Scrape or Batch action to get structured data instead of raw content:\n\n| Parser | Extracts |\n| --- | --- |\n| `@olostep/amazon-product` | Title, price, rating, reviews, images, variants |\n| `@olostep/google-search` | Result titles, URLs, snippets |\n| `@olostep/google-maps` | Business name, address, rating, reviews |\n| `@olostep/extract-emails` | Email addresses from any page |\n| `@olostep/extract-socials` | Social profile links (X, GitHub, LinkedIn, etc.) |\n| `@olostep/extract-calendars` | Google Calendar and ICS links |\n\nSee the full list in the [Olostep parser store \u2192](https://www.olostep.com/store)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#troubleshooting)  Troubleshooting\n\nAPI key rejected\n\nCopy the key directly from [olostep.com/dashboard](https://olostep.com/dashboard) with no trailing spaces. Delete and recreate the credential in n8n if the error persists.\n\nScraped content is empty\n\nIncrease **Wait Before Scraping** (try 2000\u20135000ms for JS-heavy pages). Confirm the URL is publicly accessible without a login. If a specific domain is consistently failing, contact [info@olostep.com](mailto:info@olostep.com).\n\nBatch URL format error\n\nThe **URLs to Scrape** field expects a JSON array:\n\n```\n[\\\n  { \"url\": \"https://example.com/page-1\", \"custom_id\": \"p1\" },\\\n  { \"url\": \"https://example.com/page-2\", \"custom_id\": \"p2\" }\\\n]\n```\n\nUse a Code node upstream to build this array from your data if needed.\n\nRate limit hit\n\nAdd a **Wait** node between scrape steps, or switch to **Batch Scrape URLs** instead of looping single scrapes. Check current usage in the [dashboard](https://olostep.com/dashboard).\n\nCommunity Nodes not visible in Settings\n\nOn n8n Cloud, community nodes must be enabled by a workspace owner. On self-hosted, make sure `N8N_COMMUNITY_PACKAGES_ENABLED=true` is set in your environment. See [n8n\u2019s installation guide](https://docs.n8n.io/integrations/community-nodes/installation/).\n\n* * *\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#related)  Related\n\n[**Scrapes API** \\\\\n\\\\\nFull reference for the scrape endpoint](https://docs.olostep.com/features/scrapes)\n\n[**Batches API** \\\\\n\\\\\nHow batch jobs work and how to retrieve results](https://docs.olostep.com/features/batches)\n\n[**Crawls API** \\\\\n\\\\\nCrawl configuration and result retrieval](https://docs.olostep.com/features/crawls)\n\n[**Maps API** \\\\\n\\\\\nURL discovery and filtering options](https://docs.olostep.com/features/maps)\n\n## [\u200b](https://docs.olostep.com/integrations/n8n\\#get-started)  Get Started\n\nReady to automate your web search, scraping, and crawling workflows?\n\n[**n8n Website** \\\\\n\\\\\nn8n platform](https://n8n.io/)\n\n[**Install the Node** \\\\\n\\\\\nInstall n8n-nodes-olostep and start building automated workflows](https://www.npmjs.com/package/n8n-nodes-olostep)\n\nConnect Olostep with n8n and automate your web data extraction today!\n\nWas this page helpful?\n\nYesNo",
          "metadata": {
            "og:image:height": "630",
            "og:url": "https://docs.olostep.com/integrations/n8n",
            "apple-mobile-web-app-title": "Olostep Docs",
            "viewport": "width=device-width, initial-scale=1, viewport-fit=cover",
            "twitter:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DIntegrations%26title%3DOlostep%2B%252B%2Bn8n%26description%3DAdd%2Bweb%2Bscraping%252C%2Bsearch%252C%2Bbatch%2Bjobs%252C%2Bcrawls%252C%2Band%2Bsite%2Bmaps%2Bto%2Bany%2Bn8n%2Bworkflow%252C%2Bno%2Bcode%2Brequired.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "ogDescription": "Add web scraping, search, batch jobs, crawls, and site maps to any n8n workflow, no code required.",
            "twitter:description": "Add web scraping, search, batch jobs, crawls, and site maps to any n8n workflow, no code required.",
            "generator": "Mintlify",
            "language": "en",
            "twitter:title": "Olostep + n8n - Olostep Docs",
            "robots": "index, follow",
            "og:description": "Add web scraping, search, batch jobs, crawls, and site maps to any n8n workflow, no code required.",
            "ogImage": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DIntegrations%26title%3DOlostep%2B%252B%2Bn8n%26description%3DAdd%2Bweb%2Bscraping%252C%2Bsearch%252C%2Bbatch%2Bjobs%252C%2Bcrawls%252C%2Band%2Bsite%2Bmaps%2Bto%2Bany%2Bn8n%2Bworkflow%252C%2Bno%2Bcode%2Brequired.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "og:type": "website",
            "description": "Add web scraping, search, batch jobs, crawls, and site maps to any n8n workflow, no code required.",
            "ogUrl": "https://docs.olostep.com/integrations/n8n",
            "og:title": "Olostep + n8n - Olostep Docs",
            "og:image:width": "1200",
            "twitter:card": "summary_large_image",
            "twitter:image:width": "1200",
            "twitter:image:height": "630",
            "googlebot": "index, follow",
            "ogTitle": "Olostep + n8n - Olostep Docs",
            "msapplication-TileColor": "#9563FF",
            "msapplication-config": "/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/browserconfig.xml",
            "og:site_name": "Olostep Docs",
            "title": "Olostep + n8n - Olostep Docs",
            "ogSiteName": "Olostep Docs",
            "application-name": "Olostep Docs",
            "og:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DIntegrations%26title%3DOlostep%2B%252B%2Bn8n%26description%3DAdd%2Bweb%2Bscraping%252C%2Bsearch%252C%2Bbatch%2Bjobs%252C%2Bcrawls%252C%2Band%2Bsite%2Bmaps%2Bto%2Bany%2Bn8n%2Bworkflow%252C%2Bno%2Bcode%2Brequired.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "favicon": "https://docs.olostep.com/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/android-chrome-192x192.png",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-97722032eb01",
            "sourceURL": "https://docs.olostep.com/integrations/n8n",
            "url": "https://docs.olostep.com/integrations/n8n",
            "statusCode": 200,
            "contentType": "text/html; charset=utf-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "ab26550f-2c78-4cba-8cd3-566b1367c060",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://docs.olostep.com/features/search",
          "title": "Search API - Olostep Docs",
          "description": "## [\u200b](https://docs.olostep.com/features/search#installation)  Installation\n```\npip install olostep\n```\n\n```\nnpm install olostep\n```\n\n```\n# curl is available by default on macOS, Linux, and Windows\n```\n\n## [\u200b](https://docs.olostep.com/features/search#basic-usage)  Basic usage\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.create(\"Best Answer Engine Optimization startups\")\n\nprint(search.id, len(search.links))\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n\nconst search = await client.searches.create('Best Answer Engine Optimization startups')\n\nconsole.log(search.id, search.links.length)\n```\n\n```\nimport requests, json\n\nendpoint = \"https://api.olostep.com/v1/searches\"\npayload = {\n  \"query\": \"Best Answer Engine Optimization startups\"\n}\nheaders = {\"Authorization\": \"Bearer <YOUR_API_KEY>\", \"Content-Type\": \"application/json\"}\n\nresponse = requests.post(endpoint, json=payload, headers=headers)\nprint(json.dumps(response.json(), indent=2))\n```\n\n## [\u200b](https://docs.olostep.com/features/search#retrieving-a-past-search)  Retrieving a past search\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.get(search_id=\"search_9bi0sbj9xa\")\nprint(search.id, len(search.links))\n```",
          "position": 3,
          "markdown": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/features/search#content-area)\n\nThe Olostep `/v1/searches` endpoint lets you search the web with a natural language query and get back a deduplicated list of relevant links with titles and descriptions.\n\n- Send a query in plain English\n- Get back structured links from across the web\n- Optionally scrape every returned URL in one round-trip and embed `markdown_content` / `html_content` directly into the response\n- Filter by domain, control the result count, and bound the scraping wallclock\n\nIt will search for the query semantically across the web and return results.For API details, see the [Search Endpoint API Reference](https://docs.olostep.com/api-reference/searches/create).\n\n## [\u200b](https://docs.olostep.com/features/search\\#installation)  Installation\n\nPython\n\nNode\n\ncURL\n\nNode (API)\n\nPython (API)\n\n```\npip install olostep\n```\n\n```\nnpm install olostep\n```\n\n```\n# curl is available by default on macOS, Linux, and Windows\n```\n\n```\nnpm install node-fetch\n```\n\n```\npip install requests\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#basic-usage)  Basic usage\n\nSend a natural language query and receive a list of relevant links.\n\nPython\n\nNode\n\ncURL\n\nNode (API)\n\nPython (API)\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.create(\"Best Answer Engine Optimization startups\")\n\nprint(search.id, len(search.links))\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n\nconst search = await client.searches.create('Best Answer Engine Optimization startups')\n\nconsole.log(search.id, search.links.length)\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/searches\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"query\": \"Best Answer Engine Optimization startups\"\n  }'\n```\n\n```\nconst res = await fetch('https://api.olostep.com/v1/searches', {\n  method: 'POST',\n  headers: { 'Authorization': 'Bearer <YOUR_API_KEY>', 'Content-Type': 'application/json' },\n  body: JSON.stringify({\n    query: 'Best Answer Engine Optimization startups'\n  })\n})\nconsole.log(await res.json())\n```\n\n```\nimport requests, json\n\nendpoint = \"https://api.olostep.com/v1/searches\"\npayload = {\n  \"query\": \"Best Answer Engine Optimization startups\"\n}\nheaders = {\"Authorization\": \"Bearer <YOUR_API_KEY>\", \"Content-Type\": \"application/json\"}\n\nresponse = requests.post(endpoint, json=payload, headers=headers)\nprint(json.dumps(response.json(), indent=2))\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#request-parameters)  Request parameters\n\n| Field | Type | Required | Default | Description |\n| --- | --- | --- | --- | --- |\n| `query` | string | yes | \u2014 | The search query in natural language. |\n| `limit` | integer | no | `12` | Maximum number of links to return after deduplication. Must be between `1` and `25`. |\n| `include_domains` | string\\[\\] | no | `[]` | Restrict results to these domains. Bare hosts only \u2014 leading `http(s)://` and trailing slashes are stripped automatically. |\n| `exclude_domains` | string\\[\\] | no | `[]` | Exclude results from these domains. Bare hosts only \u2014 leading `http(s)://` and trailing slashes are stripped automatically. |\n| `scrape_options` | object | no | \u2014 | When provided, every returned link is also scraped and its content embedded in the response. See [scrape\\_options](https://docs.olostep.com/features/search#scrape-options) below. |\n| `fast_mode` | boolean | no | `false` | Request a direct, low-latency search using your query verbatim. Default mode performs a broader search pass for wider result coverage. |\n\n### [\u200b](https://docs.olostep.com/features/search\\#limiting-the-number-of-results)  Limiting the number of results\n\n```\n{\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"limit\": 5\n}\n```\n\n### [\u200b](https://docs.olostep.com/features/search\\#filtering-by-domain)  Filtering by domain\n\n`include_domains` narrows results to a whitelist; `exclude_domains` filters out unwanted sources. They can be combined.\n\n```\n{\n  \"query\": \"OpenAI Sora shutdown analysis\",\n  \"include_domains\": [\"nytimes.com\", \"wsj.com\", \"bbc.com\"],\n  \"exclude_domains\": [\"pinterest.com\"]\n}\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#scrape_options)  scrape\\_options\n\nPass `scrape_options` to scrape every returned URL in parallel and embed the rendered content directly on each link. This saves a round-trip per result vs. calling `/v1/searches` and `/v1/scrapes` separately.\n\n```\n{\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"limit\": 10,\n  \"scrape_options\": {\n    \"formats\": [\"markdown\"],\n    \"remove_css_selectors\": \"default\",\n    \"timeout\": 25\n  }\n}\n```\n\n| Field | Type | Default | Description |\n| --- | --- | --- | --- |\n| `formats` | string\\[\\] | `[\"markdown\"]` | Output formats to attach to each link. For `/v1/searches`, only `\"html\"` and `\"markdown\"` are supported. Pass `[\"html\", \"markdown\"]` to receive both. |\n| `remove_css_selectors` | string | `\"default\"` | Forwarded to `/v1/scrapes`. `\"default\"` strips nav/footer/script/style/svg/dialog noise. Use `\"none\"` to disable, or pass a JSON-stringified array of selectors to remove. |\n| `timeout` | integer | `25` | Wallclock budget in **seconds** for the entire scrape phase. Must be between `1` and `60`. After this elapses, the search returns immediately \u2014 content fields will be `null` for any links that hadn\u2019t finished. |\n\n### [\u200b](https://docs.olostep.com/features/search\\#behavior)  Behavior\n\n- All links are scraped **in parallel**. The `timeout` bounds the whole batch, not each individual link.\n- Per-link scrape failures (network errors, individual page timeouts) leave that link\u2019s `markdown_content` / `html_content` as `null` while other links return normally.\n- If the global `timeout` elapses before all scrapes finish, the search responds immediately with the links it has \u2014 already-completed scrapes keep their content; in-flight ones come back with `null` content.\n- For `reddit.com/.../comments/...` URLs, the request is automatically routed through the `@olostep/reddit-post` parser and the structured JSON is rendered into clean markdown + basic HTML\n- If the combined inline content exceeds 9MB, content fields are nulled, `result.size_exceeded` is set to `true`, and you can fetch the full payload from `result.json_hosted_url`.\n\n### [\u200b](https://docs.olostep.com/features/search\\#example-with-scraping)  Example with scraping\n\nPython\n\nNode\n\ncURL\n\nNode (API)\n\nPython (API)\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.create(\n    query=\"What's going on with OpenAI's Sora shutting down?\",\n    limit=5,\n    scrape_options={\"formats\": [\"markdown\"], \"timeout\": 25},\n)\n\nfor link in search.links:\n    print(link[\"url\"], \"\u2014\", len(link.get(\"markdown_content\") or \"\"), \"chars\")\n```\n\n```\nimport Olostep, { Format } from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n\nconst search = await client.searches.create({\n  query: \"What's going on with OpenAI's Sora shutting down?\",\n  limit: 5,\n  scrapeOptions: {\n    formats: [Format.MARKDOWN],\n    timeout: 25\n  }\n})\n\nfor (const link of search.links) {\n  console.log(link.url, '\u2014', (link.markdown_content || '').length, 'chars')\n}\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/searches\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"query\": \"What'\"'\"'s going on with OpenAI'\"'\"'s Sora shutting down?\",\n    \"limit\": 5,\n    \"scrape_options\": {\n      \"formats\": [\"markdown\"],\n      \"timeout\": 25\n    }\n  }'\n```\n\n```\nconst res = await fetch('https://api.olostep.com/v1/searches', {\n  method: 'POST',\n  headers: { 'Authorization': 'Bearer <YOUR_API_KEY>', 'Content-Type': 'application/json' },\n  body: JSON.stringify({\n    query: \"What's going on with OpenAI's Sora shutting down?\",\n    limit: 5,\n    scrape_options: {\n      formats: ['markdown'],\n      timeout: 25\n    }\n  })\n})\nconst data = await res.json()\nfor (const link of data.result.links) {\n  console.log(link.url, '\u2014', (link.markdown_content || '').length, 'chars')\n}\n```\n\n```\nimport requests, json\n\nendpoint = \"https://api.olostep.com/v1/searches\"\npayload = {\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"limit\": 5,\n  \"scrape_options\": {\n    \"formats\": [\"markdown\"],\n    \"timeout\": 25\n  }\n}\nheaders = {\"Authorization\": \"Bearer <YOUR_API_KEY>\", \"Content-Type\": \"application/json\"}\n\nresponse = requests.post(endpoint, json=payload, headers=headers)\ndata = response.json()\nfor link in data[\"result\"][\"links\"]:\n  print(link[\"url\"], \"\u2014\", len(link.get(\"markdown_content\") or \"\"), \"chars\")\n```\n\n## [\u200b](https://docs.olostep.com/features/search\\#response)  Response\n\nYou will receive a `search` object in response. The `search` object contains an `id`, your original `query`, `credits_consumed`, and a `result` with a list of `links`.\n\n```\n{\n  \"id\": \"search_9bi0sbj9xa\",\n  \"object\": \"search\",\n  \"created\": 1760327323,\n  \"metadata\": {},\n  \"query\": \"What's going on with OpenAI's Sora shutting down?\",\n  \"credits_consumed\": 10,\n  \"result\": {\n    \"json_content\": \"...\",\n    \"json_hosted_url\": \"https://olostep-storage.s3.us-east-1.amazonaws.com/search_9bi0sbj9xa.json\",\n    \"size_exceeded\": false,\n    \"credits_consumed\": 10,\n    \"links\": [\\\n      {\\\n        \"url\": \"https://www.bbc.com/news/articles/c3w3e467ewqo\",\\\n        \"title\": \"OpenAI to shut down Sora video platform\",\\\n        \"description\": \"OpenAI says it will discontinue its Sora app...\",\\\n        \"markdown_content\": \"# OpenAI to shut down Sora video platform\\n\\nOpenAI says it will discontinue...\"\\\n      },\\\n      {\\\n        \"url\": \"https://www.reddit.com/r/OutOfTheLoop/comments/1s2u847/whats_going_on_with_openais_sora_shutting_down/\",\\\n        \"title\": \"What's going on with OpenAI's Sora shutting down?\",\\\n        \"description\": \"Reddit thread discussing the shutdown.\",\\\n        \"markdown_content\": \"# What's going on with OpenAI's Sora shutting down?\\n\\n*r/OutOfTheLoop \u00b7 u/rm-minus-r \u00b7 1mo ago*\\n\\n...\"\\\n      }\\\n    ]\n  }\n}\n```\n\nEach link in `result.links` contains:\n\n| Field | Type | Description |\n| --- | --- | --- |\n| `url` | string | The URL of the search result. |\n| `title` | string | The title of the result page. |\n| `description` | string | A short snippet describing the result. |\n| `markdown_content` | string | Markdown content of the page. Only present when `scrape_options.formats` includes `\"markdown\"`. `null` if the scrape failed, was empty, or hit the global timeout. |\n| `html_content` | string | HTML content of the page. Only present when `scrape_options.formats` includes `\"html\"`. `null` on failure/timeout. |\n\nThe full result is also available as a hosted JSON file at `result.json_hosted_url` \u2014 useful when `result.size_exceeded` is `true`.\n\n## [\u200b](https://docs.olostep.com/features/search\\#retrieving-a-past-search)  Retrieving a past search\n\n`GET /v1/searches/{search_id}` returns whatever was persisted at search time, including any scraped content. It\u2019s a pure idempotent read \u2014 no re-scraping, no re-billing. Older searches without `scrape_options` simply have no per-link content fields.\n\nPython\n\nNode\n\ncURL\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\n\nsearch = client.searches.get(search_id=\"search_9bi0sbj9xa\")\nprint(search.id, len(search.links))\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\n\nconst search = await client.searches.get('search_9bi0sbj9xa')\nconsole.log(search.id, search.links.length)\n```\n\n```\ncurl -s \"https://api.olostep.com/v1/searches/search_9bi0sbj9xa\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\"\n```\n\nSee [Get Search](https://docs.olostep.com/api-reference/searches/get) for full details.\n\n## [\u200b](https://docs.olostep.com/features/search\\#pricing)  Pricing\n\nEach search costs **5 credits** for the search itself.When `scrape_options` is provided, each scraped page is billed at the standard `/v1/scrapes` rate (typically 1 credit per page; some parsers cost more). The total is returned in `credits_consumed`.Examples:\n\n| Request | `credits_consumed` |\n| --- | --- |\n| Search only | `5` |\n| Search + 5 scraped pages (1 credit each) | `10` |\n\nWas this page helpful?\n\nYesNo",
          "metadata": {
            "apple-mobile-web-app-title": "Olostep Docs",
            "ogUrl": "https://docs.olostep.com/features/search",
            "msapplication-TileColor": "#9563FF",
            "og:url": "https://docs.olostep.com/features/search",
            "og:image:width": "1200",
            "title": "Search API - Olostep Docs",
            "twitter:image:width": "1200",
            "language": "en",
            "ogImage": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DFeatures%26title%3DSearch%2BAPI%26description%3DSearch%2Bthe%2Bweb%2Bwith%2Bnatural%2Blanguage%2Band%2Bget%2Bstructured%2Blinks%2B%25E2%2580%2594%2BOlostep%2527s%2BAI%2Bsearch%2BAPI.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "application-name": "Olostep Docs",
            "og:type": "website",
            "ogDescription": "Search the web with natural language and get structured links \u2014 Olostep's AI search API.",
            "og:description": "Search the web with natural language and get structured links \u2014 Olostep's AI search API.",
            "og:site_name": "Olostep Docs",
            "ogTitle": "Search API - Olostep Docs",
            "description": "Search the web with natural language and get structured links \u2014 Olostep's AI search API.",
            "generator": "Mintlify",
            "twitter:title": "Search API - Olostep Docs",
            "twitter:card": "summary_large_image",
            "googlebot": "index, follow",
            "twitter:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DFeatures%26title%3DSearch%2BAPI%26description%3DSearch%2Bthe%2Bweb%2Bwith%2Bnatural%2Blanguage%2Band%2Bget%2Bstructured%2Blinks%2B%25E2%2580%2594%2BOlostep%2527s%2BAI%2Bsearch%2BAPI.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "twitter:image:height": "630",
            "msapplication-config": "/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/browserconfig.xml",
            "og:image:height": "630",
            "ogSiteName": "Olostep Docs",
            "robots": "index, follow",
            "twitter:description": "Search the web with natural language and get structured links \u2014 Olostep's AI search API.",
            "og:title": "Search API - Olostep Docs",
            "og:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DFeatures%26title%3DSearch%2BAPI%26description%3DSearch%2Bthe%2Bweb%2Bwith%2Bnatural%2Blanguage%2Band%2Bget%2Bstructured%2Blinks%2B%25E2%2580%2594%2BOlostep%2527s%2BAI%2Bsearch%2BAPI.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "viewport": "width=device-width, initial-scale=1, viewport-fit=cover",
            "favicon": "https://docs.olostep.com/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/android-chrome-192x192.png",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-9b78a07e3ff9",
            "sourceURL": "https://docs.olostep.com/features/search",
            "url": "https://docs.olostep.com/features/search",
            "statusCode": 200,
            "contentType": "text/html; charset=utf-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "bb60d55e-833d-449f-a4a6-73711a2ac7b5",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages",
          "title": "Estimating storage from the number of index pages - IBM",
          "description": "# Estimating storage from the number of index pages\n## Procedure [Copy to clipboard](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages#taskdb2z_estimatestoragefromindex__steps__1)\n19. Calculate the total index pages total index pages is approximately MAX(4, (1 + tree pages + free pages + space map pages))",
          "position": 4,
          "markdown": "[Skip to content](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages#main-content)[IBM logo](https://www.ibm.com/)Documentation\n[Announcements](https://www.ibm.com/docs/en/announcements)[IBM Redbooks](https://www.ibm.com/docs/en/redbooks)[IBM Product Documentation Directory](https://www.ibm.com/docs/en/products)Table of contentsDark mode\n\nYour current region:\n\n**Unknown \u2013 Unknown**\n\nNo other regions available\n\nMy IBM\nLog in\n\n\nDark mode\n\nDb2 for z/OS\n\nClose table of contents\n\nChange version\n\nSelect\n\n13.0.012.0.011.0.0\n\nShow full table of contents\n\nFilter on titles\n\n- [Welcome to Db2 13 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=welcome-db2-13-zos)\n\n- [About Db2 13 documentation](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=about-db2-13-documentation)\n\n- [Getting started with Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=getting-started-db2-zos)\n\n- [What's new in Db2 13](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=whats-new-in-db2-13)\n\n- [Adopting new capabilities in Db2 13 continuous delivery](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=adopting-new-capabilities-in-db2-13-continuous-delivery)\n\n- [Installing or migrating to Db2 13](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=installing-migrating-db2-13)\n\n- [Db2 application DevOps solutions](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-application-devops-solutions)\n\n- [Administering Db2](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=administering-db2)\n\n\n\n\n\n\n\n  - [Designing and implementing Db2 databases](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-designing-implementing-databases)\n\n\n\n\n\n\n\n    - [Database objects and relationships](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=databases-database-objects-relationships)\n\n    - [Implementing your database design](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=databases-implementing-your-database-design)\n\n\n\n\n\n\n\n      - [Implementing databases](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-databases)\n\n      - [Implementing storage groups](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-storage-groups)\n\n      - [Implementing table spaces](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-table-spaces)\n\n      - [Implementing tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-tables)\n\n      - [Implementing views](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-views)\n\n      - [Implementing indexes](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-indexes)\n\n      - [Implementing schemas](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-schemas)\n\n      - [Loading data into tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-loading-data-into-tables)\n\n      - [implementing stored procedures](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-stored-procedures)\n\n      - [Implementing relationships with referential constraints](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-relationships-referential-constraints)\n\n      - [implementing triggers](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-triggers)\n\n      - [Implementing user-defined functions](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-user-defined-functions)\n\n      - [Implementing Db2 system-defined routines](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-implementing-db2-system-defined-routines)\n\n      - [Obfuscating source code of SQL procedures, SQL functions, and triggers](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=iydd-obfuscating-source-code-sql-procedures-sql-functions-triggers)\n\n      - [Estimating disk storage for user data](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-estimating-disk-storage-user-data)\n\n\n\n\n\n\n\n        - [General approach to estimating storage](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=data-general-approach-estimating-storage)\n\n        - [Calculating the space required for a table](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=data-calculating-space-required-table)\n\n        - [Calculating the space required for an index](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=data-calculating-space-required-index)\n\n\n\n\n\n\n\n          - [Levels of index pages](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-levels-pages)\n\n          - [Estimating storage from the number of index pages](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages)\n\n\n      - [Identifying databases that might exceed the OBID limit](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=design-identifying-databases-that-might-exceed-obid-limit)\n\n\n    - [Altering your database design](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=databases-altering-your-database-design)\n\n\n  - [Operation and recovery](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-operation-recovery)\n\n  - [Exit routines](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-exit-routines)\n\n  - [Data sharing](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-data-sharing)\n\n  - [International data](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-international-data)\n\n  - [Db2 REST services](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-rest-services)\n\n  - [IBM Text Search for Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-text-search-zos)\n\n  - [IBM Spatial Support for Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-spatial-support-zos)\n\n  - [IBM SQL Tuning Services](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-sql-tuning-services)\n\n  - [IBM Db2 for z/OS Agent](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-zos-agent)\n\n\n- [Programming for Db2 for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=programming-db2-zos)\n\n- [Securing Db2](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=securing-db2)\n\n- [Managing Db2 performance](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=managing-db2-performance)\n\n- [Running AI queries with SQL Data Insights](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=running-ai-queries-sql-data-insights)\n\n- [Enabling Db2 for IBM Db2 Analytics Accelerator for z/OS](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=enabling-db2-db2-analytics-accelerator-zos)\n\n- [Implementing Db2 stored procedures](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=implementing-db2-stored-procedures)\n\n- [Db2 SQL](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-sql)\n\n- [Db2 commands](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-commands)\n\n- [Db2 Utilities](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-utilities)\n\n- [Db2 catalog tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-catalog-tables)\n\n- [Db2-supplied user tables](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-supplied-user-tables)\n\n- [Db2 messages](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-messages)\n\n- [Db2 codes](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-codes)\n\n- [IRLM messages and codes](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=irlm-messages-codes)\n\n- [Troubleshooting problems in Db2](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=troubleshooting-problems-in-db2)\n\n- [Db2 glossary](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=db2-glossary)\n\n- [Notices](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=notices)\n\n[Announcements & sales manuals](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?announcement=all)\n\n[Download PDF](https://www.ibm.com/docs/en/SSEPEK_13.0.0/home/src/tpc/db2z_pdfmanuals.html \"Download PDF\")\n\nOffline docs\n\nFocus sentinel\n\n## Security update required for IBM Docs Offline\n\nClose\n\nA critical security vulnerability has been identified in IBM Docs Offline. You must download and install the latest version to stay protected.\n\nVisit the [Docs Offline page](https://www.ibm.com/docs/en/offline) to download the updated version, or [view the security bulletin](https://www.ibm.com/support/pages/node/7283484 \"Opens in a new tab\") for full details.\n\nGo to Docs Offline pageContinue download\n\nFocus sentinel\n\n[Get hands-on experience with IBM tech\u00a0\u00a0Join one of the largest technical IBM community gatherings!\u00a0\u00a0\u2192](https://www.ibm.com/events/techxchange)\n\nChange version\n\n13.0.012.0.011.0.0\n\nWas this topic helpful?\n\npositive feedback\n\nnegative feedback\n\nFocus sentinel\n\n## Rate this content\n\nClose\n\nGreat! Let us know what you found helpful.\n\nComment 0/750\n\nNote: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support.\n\nCancelSubmit\n\nFocus sentinel\n\nFocus sentinel\n\n## Rate this content\n\nClose\n\nWhat can we do to improve the content?\n\nComment 0/750\n\nNote: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support.\n\nCancelSubmit\n\nFocus sentinel\n\nFocus sentinel\n\n## Provide more feedback\n\nClose\n\nComment 0/750\n\nNote: This feedback goes to your product's documentation team and does not include a response. Issues that require a response should go through IBM support.\n\nCancelSubmit\n\nFocus sentinel\n\nFocus sentinel\n\nClose\n\nRate this content\n\nThank you for your feedback!\n\nTogether, we can continue to improve IBM Documentation.\n\nReturn to topic\n\nFocus sentinel\n\nFocus sentinel\n\nClose\n\n## Thank you for your submission.\n\nSubmissions are limited to 1 per day per topic.\n\nFocus sentinel\n\nFocus sentinel\n\nClose\n\n## Error submitting rating\n\nThere has been an error sending your feedback to the team. Your comment was saved locally, if not in an incognito browser, and will be available when attempting to submit feedback again.\n\nPlease try again later.\n\nFocus sentinel\n\n# Estimating storage from the number of index pages\n\nLast Updated: 2026-01-07\n\nBefore you run a LOAD utility job to load an index, estimate\nthe future storage requirements of the index.\n\n## About this task [Copy to clipboard](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages\\#taskdb2z_estimatestoragefromindex__context__1 \"Copy to clipboard\")\n\nAn index key on an auxiliary table for LOBs is 19 bytes and uses the same formula as other indexes. The RID value that is stored within the index is 4, 5, or 7 bytes, depending on the table space type.\n\nIn general, the\nlength of the index key is the sum of the lengths of all the columns\nof the key, plus the number of columns that allow nulls. The length\nof a varying-length column is the maximum length if the index is padded.\nOtherwise, if an index is not padded, estimate the length of a varying-length\ncolumn to be the average length of the column data, and add a two-byte\nlength field to the estimate. You can retrieve the value of the AVGKEYLEN\ncolumn in the SYSIBM.SYSINDEXES catalog table to determine the average\nlength of keys within an index.\n\nThe following index calculations\nare intended only to help you estimate the storage required for an\nindex. Because there is no way to predict the exact number of duplicate\nkeys that can occur in an index, the results of these calculations\nare not absolute. It is possible, for example, that for a nonunique\nindex, more index entries than the calculations indicate might be\nable to fit on an index page.\n\nImportant: Space allocation\nparameters are specified in kilobytes.\n\nIn the following calculations,\nassume the following:\n\nkThe length of the index key.nThe average number of data records per distinct key value of a\nnonunique index. For example:\n\n- a = number of data records per index\n- b = number of distinct key values per index\n- n = a / b\n\nfThe value of PCTFREE.pThe value of FREEPAGE.rThe record identifier (RID) length. Use 4, 5, or 7 bytes, depending on the table space type:\n\n| Table space type | r value to use |\n| --- | --- |\n| Partition-by-growth (PBG UTS) | 5 |\n| Partition-by-range (PBR UTS) with relative page numbering | 7 |\n| Partition-by-range (PBR UTS) with absolute page numbering | 5 |\n| Non-UTS defined with `DSSIZE 4G` or greater | 5 |\n| Non-UTS defined with `LARGE` | 5 |\n| Other non-UTS types | 4 |\n\nSThe value of the page size minus the length of the page header\nand page tail.FLOORThe operation of discarding the decimal portion of a real number.CEILINGThe operation of rounding a real number up to the next highest\ninteger.MAXThe operation of selecting the highest integer value.\n\n## Procedure [Copy to clipboard](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages\\#taskdb2z_estimatestoragefromindex__steps__1 \"Copy to clipboard\")\n\nTo estimate index storage size, complete the following calculations:\n\n1. Calculate the pages for a unique index.\n1. Calculate the total leaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately k \\+ r \\+ 3\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR((100 - f) \u00d7 S / 100)\n\n      3. Calculate the entries per page\n         entries per page is approximately FLOOR(usable space per page / space per key)\n\n      4. Calculate the **total leaf pages**\n         total leaf pages is approximately CEILING(number of table rows / entries per page)\n\n\n2. Calculate the total nonleaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately k \\+ 7\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR(MAX(90, (100 - f )) \u00d7 S /100)\n\n      3. Calculate the entries per page\n         entries per page is approximately FLOOR(usable space per page / space per key)\n\n      4. Calculate the minimum child pages\n         minimum child pages is approximately MAX(2, (entries per page + 1))\n\n      5. Calculate the level 2 pages\n         level 2 pages is approximately CEILING(total leaf pages / minimum child pages)\n\n      6. Calculate the level 3 pages\n         level 3 pages is approximately CEILING(level 2 pages / minimum child pages)\n\n      7. Calculate the level x pages\n         level x pages is approximately CEILING(previous level pages / minimum child pages)\n\n      8. Calculate the **total nonleaf pages**\n         total nonleaf pages is approximately (level 2 pages \\+ level 3 pages \\+ ... \\+ level x pages until the number of level x pages = 1)\n2. Calculate the pages for a nonunique index.\n1. Calculate the total leaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately 4 + k \\+ (n \u00d7 (r+1))\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR((100 - f ) \u00d7 S / 100)\n\n      3. Calculate the key entries per page\n         key entries per page is approximately n\u00d7 (usable space per page / space per key)\n\n      4. Calculate the remaining space per page\n         remaining space per page is approximately usable space per page \\- (key entries per page / n) \u00d7space per key\n\n      5. Calculate the data records per partial entry\n         data records per partial entry is approximately FLOOR((remaining space per page \\- (k \\+ 4)) / 5)\n\n      6. Calculate the partial entries per page\n         partial entries per page is approximately (n / CEILING(n / data records per partial entry)) if data records per partial entry >= 1, or 0 if data records per partial entry < 1\n\n      7. Calculate the entries per page\n         entries per page is approximately MAX(1, (key entries per page \\+ partial entries per page))\n\n      8. Calculate the **total leaf pages**\n         total leaf pages is approximately CEILING(number of table rows / entries per page)\n\n\n2. Calculate the total nonleaf pages\n\n\n      1. Calculate the space per key\n         space per key is approximately k \\+ r \\+ 7\n\n      2. Calculate the usable space per page\n         usable space per page is approximately FLOOR (MAX(90, (100- f))\u00d7 S / 100)\n\n      3. Calculate the entries per page\n         entries per page is approximately FLOOR((usable space per page / space per key)\n\n      4. Calculate the minimum child pages\n         minimum child pages is approximately MAX(2, (entries per page \\+ 1))\n\n      5. Calculate the level 2 pages\n         level 2 pages is approximately CEILING(total leaf pages / minimum child pages)\n\n      6. Calculate the level 3 pages\n         level 3 pages is approximately CEILING(level 2 pages / minimum child pages)\n\n      7. Calculate the level x pages\n         level x pages is approximately CEILING(previous level pages / minimum child pages)\n\n      8. Calculate the **total nonleaf pages**\n         total nonleaf pages is approximately (level 2 pages \\+ level 3 pages \\+ ... \\+ level x pages until x = 1)\n3. Calculate the pages for an index that is not compressed.\n1. Calculate the usable space per leaf page:\n\n\n      usable space per leaf page is approximately FLOOR((100 - f) \u00d7 S / 100)\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 62 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n2. Calculate the usable space per nonleaf page:\n\n\n      usable space per nonleaf page is approximately FLOOR (MAX (90, (100 - f ) ) \u00d7 S / 100)\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 48 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n3. Calculate the usable space per space map:\n\n\n      usable space per space map is approximately CEILING ( (tree pages \\+ free pages) / S), where S equals (page size \u2212 header length \u2212 tail length) \u00d7 2 \u2212 1.\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 28 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n4. Calculate the pages for a compressed index.\n1. Calculate the usable space per leaf page:\n\n\n      usable space per leaf page is approximately FLOOR((100 - f) \u00d7 S / 100)\n\n\n\n      The page size can be 4096 bytes (4 KB), 8192 bytes (8 KB), 16384 bytes (16 KB), or 32768 bytes (32 KB). The length of the page header is 66 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n2. Calculate the usable space per nonleaf page:\n\n\n      usable space per nonleaf page is approximately FLOOR (MAX (90, (100 - f ) ) \u00d7 S / 100)\n\n\n\n      The page size is 4096 bytes for 4 KB, 8 KB, 16 KB, and 32 KB page sizes. The length of the page header is 48 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n\n3. Calculate the usable space per space map:\n\n\n      usable space per space map is approximately CEILING ( (tree pages + free pages) / S), where S equals (page size \u2212 header length \u2212 tail length) \u00d7 2 \u2212 1.\n\n\n\n      The page size is 4096 bytes for 4 KB, 8 KB, 16 KB, and 32 KB page sizes. The length of the page header is 28 bytes. The length of the page tail is 20 bytes for 10-byte RBA or LRSN format, or 2 bytes for 6-byte RBA or LRSN format.\n5. Calculate the total space requirement by estimating the number of kilobytes required for an index built by the LOAD utility.\n1. Calculate the free pages\n\n\n      free pages is approximately FLOOR(total leaf pages / p), or 0 if p = 0\n\n2. Calculate the space map pages\n\n\n      space map pages is approximately CEILING((tree pages \\+ free pages) / S)\n\n3. Calculate the tree pages\n\n\n      tree pages is approximately MAX(2, (total leaf pages \\+ total nonleaf pages))\n\n4. Calculate the total index pages\n\n\n      total index pages is approximately MAX(4, (1 + tree pages \\+ free pages \\+ space map pages))\n\n5. Calculate the **total space requirement**\n\n\n      total space requirement is approximately 4 \u00d7 (total index pages \\+ 2)\n\n## Example [Copy to clipboard](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages\\#taskdb2z_estimatestoragefromindex__example__1 \"Copy to clipboard\")\n\nIn the following example of the entire calculation, assume that an index is defined with these characteristics:\n\n- The index is unique.\n- The table it indexes has 100000 rows.\n- The key is a single column defined as CHAR(10) NOT NULL.\n- The value of PCTFREE is 5.\n- The value of FREEPAGE is 4.\n- The page size is 4 KB.\n\n| Quantity | Calculation | Result |\n| --- | --- | --- |\n| Length of key<br>Average number of duplicate keys<br>PCTFREE<br>FREEPAGE | k<br>n<br>f<br>p | 10<br>1<br>5<br>4 |\n| **Calculate total leaf pages**<br>Space per key<br>Usable space per page<br>Entries per page<br>Total leaf pages | k \\+ 7<br>FLOOR((100 - f ) \u00d7 4032/100)<br>FLOOR(usable space per page / space per key)<br>CEILING(number of table rows / entries per page) | 17<br>3844<br>225<br>445 |\n| **Calculate total nonleaf pages**<br>Space per key<br>Usable space per page<br>Entries per page<br>Minimum child pages<br>Level 2 pages<br>Level 3 pages<br>Total nonleaf pages | k \\+ 7<br>FLOOR(MAX(90, (100 - f )) \u00d7 4046/100)<br>FLOOR(usable space per page / space per key)<br>MAX(2, (entries per page \\+ 1))<br>CEILING(total leaf pages / minimum child pages)<br>CEILING(level 2 pages / minimum child pages)<br>(level 2 pages \\+ level 3 pages \\+\u2026\\+ level x pages until x = 1) | 17<br>3843<br>226<br>227<br>2<br>1<br>3 |\n| **Calculate total space required**<br>Free pages<br>Tree pages<br>Space map pages<br>Total index pages<br>TOTAL SPACE REQUIRED, in KB | FLOOR(total leaf pages / p), or 0 if p = 0<br>MAX(2, (total leaf pages \\+ total nonleaf pages))<br>CEILING((tree pages \\+ free pages)/8131)<br>MAX(4, (1 + tree pages \\+ free pages \\+ space map pages))<br>4 \u00d7 (total index pages \\+ 2) | 111<br>448<br>1<br>561<br>2252 |\n\nTable 1. Sample of the total space requirement for a unique index\n\nWhile IBM values the use of inclusive language, terms that are outside of IBM's direct influence, for the sake of maintaining user understanding, are sometimes required. As other industry leaders join IBM in embracing the use of inclusive language, IBM will continue to update the documentation to reflect those changes.\n\n\u00a9 Copyright IBM Corporation 1983, 2026\n\n[IBM logo](https://www.ibm.com/)\n\nArabic / \u0639\u0631\u0628\u064a\u0629\n\nBulgarian / \u0411\u044a\u043b\u0433\u0430\u0440\u0441\u043a\u0438\n\nCatalan / Catal\u00e0\n\nCzech / \u010ce\u0161tina\n\nDanish / Dansk\n\nGerman / Deutsch\n\nGreek / \u0395\u03bb\u03bb\u03b7\u03bd\u03b9\u03ba\u03ac\n\nEnglish\n\nSpanish / Espa\u00f1ol\n\nFinnish / Suomi\n\nFrench / Fran\u00e7ais\n\nCroatian / Hrvatski\n\nHungarian / Magyar\n\nItalian / Italien\n\nHebrew / \u05e2\u05d1\u05e8\u05d9\u05ea\n\nJapanese / \u65e5\u672c\u8a9e\n\nKorean / \ud55c\uad6d\uc5b4\n\nKazakh / \u049a\u0430\u0437\u0430\u049b\u0448\u0430\n\nDutch / Nederlands\n\nNorwegian / Norsk\n\nPolish / polski\n\nPortuguese/Brazil / Portugu\u00eas/Brasil\n\nPortuguese/Portugal / Portugu\u00eas/Portugal\n\nRomanian / Rom\u00e2n\u0103\n\nRussian / \u0420\u0443\u0441\u0441\u043a\u0438\u0439\n\nSlovak / Sloven\u010dina\n\nSlovenian / sloven\u0161\u010dina\n\nSerbian / srpski\n\nSwedish / Svenska\n\nThai / \u0e20\u0e32\u0e29\u0e32\u0e44\u0e17\u0e22\n\nTurkish / T\u00fcrk\u00e7e\n\nVietnamese / Vi\u00ea\u0323t\n\nChinese Simplified / \u7b80\u4f53\u4e2d\u6587\n\nChinese Traditional / \u7e41\u9ad4\u4e2d\u6587[IBM logo](https://www.ibm.com/)Contact IBMPrivacyTerms of useAccessibilityCookie Preferences\n\nArabic / \u0639\u0631\u0628\u064a\u0629\n\nBulgarian / \u0411\u044a\u043b\u0433\u0430\u0440\u0441\u043a\u0438\n\nCatalan / Catal\u00e0\n\nCzech / \u010ce\u0161tina\n\nDanish / Dansk\n\nGerman / Deutsch\n\nGreek / \u0395\u03bb\u03bb\u03b7\u03bd\u03b9\u03ba\u03ac\n\nEnglish\n\nSpanish / Espa\u00f1ol\n\nFinnish / Suomi\n\nFrench / Fran\u00e7ais\n\nCroatian / Hrvatski\n\nHungarian / Magyar\n\nItalian / Italien\n\nHebrew / \u05e2\u05d1\u05e8\u05d9\u05ea\n\nJapanese / \u65e5\u672c\u8a9e\n\nKorean / \ud55c\uad6d\uc5b4\n\nKazakh / \u049a\u0430\u0437\u0430\u049b\u0448\u0430\n\nDutch / Nederlands\n\nNorwegian / Norsk\n\nPolish / polski\n\nPortuguese/Brazil / Portugu\u00eas/Brasil\n\nPortuguese/Portugal / Portugu\u00eas/Portugal\n\nRomanian / Rom\u00e2n\u0103\n\nRussian / \u0420\u0443\u0441\u0441\u043a\u0438\u0439\n\nSlovak / Sloven\u010dina\n\nSlovenian / sloven\u0161\u010dina\n\nSerbian / srpski\n\nSwedish / Svenska\n\nThai / \u0e20\u0e32\u0e29\u0e32\u0e44\u0e17\u0e22\n\nTurkish / T\u00fcrk\u00e7e\n\nVietnamese / Vi\u00ea\u0323t\n\nChinese Simplified / \u7b80\u4f53\u4e2d\u6587\n\nChinese Traditional / \u7e41\u9ad4\u4e2d\u6587\n\n[![close icon](https://consent.trustarc.com/get?name=ibm_close_icon.svg)](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages#)\n\nIBM web domains\n\nibm.com, ibm.org, ibm-zcouncil.com, insights-on-business.com, jazz.net, mobilebusinessinsights.com, promontory.com, proveit.com, ptech.org, s81c.com, securityintelligence.com, skillsbuild.org, softlayer.com, storagecommunity.org, think-exchange.com, thoughtsoncloud.com, alphaevents.webcasts.com, ibm-cloud.github.io, ibmbigdatahub.com, bluemix.net, mybluemix.net, ibm.net, ibmcloud.com, galasa.dev, blueworkslive.com, swiss-quantum.ch, blueworkslive.com, cloudant.com, ibm.ie, ibm.fr, ibm.com.br, ibm.co, ibm.ca, community.watsonanalytics.com, datapower.com, skills.yourlearning.ibm.com, bluewolf.com, carbondesignsystem.com, openliberty.io\n\n![close icon](https://consent.trustarc.com/get?name=ibm_close_icon.svg)\n\nAbout cookies on this siteOur websites require some cookies to function properly (required). In addition, other cookies may be used with your consent to analyze site usage, improve the user experience and for advertising.For more information, please review your cookie\u00a0preferences\u00a0options. By visiting our website, you agree to our processing of information as described in IBM\u2019s [privacy\u00a0statement](https://www.ibm.com/privacy).\u00a0 To provide a smooth navigation, your cookie preferences will be shared across the IBM web domains listed\u00a0[here](https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages#truste_domain_list).\n\nAccept AllMore options\n\n![](https://analytics.o11.tech/e/a.gif?aqet=pv&evid=e3ba4354-61ac-487e-a753-01d0414b46d7&aq_m=1&pubid=b0a6ae0081d4f8392fc41064e9ae7fb9&dmn=www.ibm.com&tt=tcs.dhj&cid=c076&lbl=c076&flbl=pxcel&ll=e&ver=1.2115.265&ell=e&cck=_autid&pn=%2Fdocs%2Fen%2Fdb2-for-zos%2F13.0.0&qs=topic%3Dindex-estimating-storage-from-number-pages&rdn=www.google.com&rpn=%2F&rqs=na&cc=US&cont=NA&rc=VA&urls=!1!0!b-13w,!1!0!b-13x,!1!0!b-13y,!1!0!b-144,!1!0!b-14i&rnd=1788514043583&cid=c076&version=1.2115.265&cc=US&cont=NA&repeat=0&htmLcy=136)\n\n![](https://a.usbrowserspeed.com/cs?pid=8046b079bd55b4d00035e43c6d9f73c385a7184438a37f8e4401bf0b831c02f6&puid=4908fe26891c9b4c6351c2cee71801d1;e3ba4354-61ac-487e-a753-01d0414b46d7;6a9a8efb8e6bdc174b1d04b8&r=https%3A%2F%2Fanalytics.o11.tech%2F1%2Fe%2Fc.gif%3Faqet%3Didsync%26puu%3D%24%7BDEVICE_ID%7D)\n\n![](https://analytics.o11.tech/e/a.gif?aqet=pv&evid=e3ba4354-61ac-487e-a753-01d0414b46d7&aq_m=1&pubid=4908fe26891c9b4c6351c2cee71801d1&dmn=www.ibm.com&tt=tcs.dhj&cid=c076&lbl=c076&flbl=pxcel&ll=e&ver=1.2115.265&ell=e&cck=_autid&pn=%2Fdocs%2Fen%2Fdb2-for-zos%2F13.0.0&qs=topic%3Dindex-estimating-storage-from-number-pages&rdn=www.google.com&rpn=%2F&rqs=na&cc=US&cont=NA&rc=VA&urls=!1!0!b-13w,!1!0!b-13x,!1!0!b-13y,!1!0!b-144,!1!0!b-14i&rnd=1788514043563&cid=c076&version=1.2115.265&cc=US&cont=NA&repeat=0&htmLcy=110)\n\n![](https://a.usbrowserspeed.com/cs?pid=8046b079bd55b4d00035e43c6d9f73c385a7184438a37f8e4401bf0b831c02f6&puid=32c5475300e9fc6db3d3688149e141ea;e3ba4354-61ac-487e-a753-01d0414b46d7;6a9a8efb8e6bdc174b1d04b8&r=https%3A%2F%2Fanalytics.o11.tech%2F1%2Fe%2Fc.gif%3Faqet%3Didsync%26puu%3D%24%7BDEVICE_ID%7D)\n\n![](https://analytics.o11.tech/e/a.gif?aqet=pv&evid=e3ba4354-61ac-487e-a753-01d0414b46d7&aq_m=1&pubid=32c5475300e9fc6db3d3688149e141ea&dmn=www.ibm.com&tt=tcs.dhj&cid=c076&lbl=c076&flbl=pxcel&ll=e&ver=1.2115.265&ell=e&cck=_autid&pn=%2Fdocs%2Fen%2Fdb2-for-zos%2F13.0.0&qs=topic%3Dindex-estimating-storage-from-number-pages&rdn=www.google.com&rpn=%2F&rqs=na&cc=US&cont=NA&rc=VA&urls=!1!0!b-13w,!1!0!b-13x,!1!0!b-13y,!1!0!b-144,!1!0!b-14i&rnd=1788514043600&cid=c076&version=1.2115.265&cc=US&cont=NA&repeat=0&htmLcy=159)",
          "metadata": {
            "ogTitle": "Db2 for z/OS",
            "og:url": "https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages",
            "og:title": "Db2 for z/OS",
            "ogLocale": "en-US",
            "geo.country": "US",
            "robots": "index,follow",
            "dcterms.date": "2026-01-07",
            "language": "en-US",
            "ogDescription": "Before you run a LOAD utility job to load an index, estimate the future storage requirements of the index.",
            "og:description": "Before you run a LOAD utility job to load an index, estimate the future storage requirements of the index.",
            "title": "Estimating storage from the number of index pages - IBM Documentation",
            "ogUrl": "https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages",
            "description": "Before you run a LOAD utility job to load an index, estimate the future storage requirements of the index.",
            "publishedTime": "2026-01-07",
            "og:published_time": "2026-01-07",
            "article:published_time": "2026-01-07",
            "ogImage": "https://1.www.s81c.com/common/images/ibm-leadspace-1200x627.jpg",
            "keywords": "indexes, storage, estimating",
            "dcterms.rights": "\u00a9 Copyright IBM Corporation 2026",
            "viewport": "width=device-width,initial-scale=1",
            "og:image": "https://1.www.s81c.com/common/images/ibm-leadspace-1200x627.jpg",
            "og:type": "website",
            "og:locale": "en-US",
            "canonical": "https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages",
            "favicon": "https://www.ibm.com/favicon.ico",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-9fd0773a6fff",
            "sourceURL": "https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages",
            "url": "https://www.ibm.com/docs/en/db2-for-zos/13.0.0?topic=index-estimating-storage-from-number-pages",
            "statusCode": 200,
            "contentType": "text/html; charset=utf-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "4a454fff-1433-4409-bded-90aa4cc7ac28",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918",
          "title": "Fetching all docs in an app search index - Elastic Discuss",
          "description": "# [Fetching all docs in an app search index](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918)\n## post by frankjoh2 on Dec 7, 2018\n_\"total_pages\": 92,_\n\n_\"size\": 1000_\n\nWorking as expected and showing all 91035 docs in the index and that there are a total of 92 pages.\n\n_\"total_pages\": 0,_\n\n_\"size\": 1000_\n\n## post by goodroot on Dec 7, 2018\n```rust\n\ncurl -X GET 'https://host-xxxxxx.api.swiftype.com/api/as/v1/engines/example-engine/documents/list' \\\n-H 'Content-Type: application/json' \\\n-H 'Authorization: Bearer private-xxxxxxxxxxxxxxxxxxxx' \\\n-d '{\n  \"page\": {\n    \"current\": 1,\n    \"size\": 100\n  }\n}'\n```",
          "position": 5,
          "markdown": "[Skip to where you left off (last reply, post 6)](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/6) [Skip to top](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/1)\n\n[Skip to main content](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918#main-container)\n\n# [Fetching all docs in an app search index](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918)\n\n[Elastic Search](https://discuss.elastic.co/c/search/84)\n\n- [elastic-app-search](https://discuss.elastic.co/tag/elastic-app-search/118)\n\nYou have selected **0** posts.\n\n[select all](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918)\n\n[cancel selecting](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918)\n\n2.1k\nviews\n\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)3](https://discuss.elastic.co/u/frankjoh2 \"frankjoh2\")\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)2](https://discuss.elastic.co/u/goodroot \"goodroot\")\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/1 \"Jump to the first post\")\n\n6 / 6\n\n\nJan 2019\n\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/6)\n\n## post by frankjoh2 on Dec 7, 2018\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)](https://discuss.elastic.co/u/frankjoh2)\n\n[frankjoh2](https://discuss.elastic.co/u/frankjoh2)\n\n1\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918 \"Post date\")\n\nHi,\n\nI'm trying to get the ids of all documents in my app search index. I tried to simply iteration of searches with an empty query and incrementally increasing the page-number. This seemed to work fine for the first 10 requests.\n\nThis is what my requests looks like:\n\n{\n\n\"query\": \"\",\n\n\"result\\_fields\": {\n\n\"id\": { \"raw\": {} }\n\n},\n\n\"page\": {\"size\": 1000,\"current\": **\\[PAGENUMER\\]** }\n\n}\n\nWhere **\\[PAGENUMBER\\]** is 1 for the first request, 2 for second and so on\u2026\n\nThis is the result of the 10th request:\n\n_{_\n\n_\"meta\": {_\n\n_\"warnings\": \\[\\],_\n\n_\"page\": {_\n\n_\"current\": 10,_\n\n_\"total\\_pages\": 92,_\n\n_\"total\\_results\": 91035,_\n\n_\"size\": 1000_\n\n_},_\n\n_\"request\\_id\": \"39684c716fe14725a70406a1a71789e4\"_\n\n_},_\n\n_\"results\": \\[...\\] <--- 1000 results here_\n\n_}_\n\nWorking as expected and showing all 91035 docs in the index and that there are a total of 92 pages.\n\nBut the result of the 11th request:\n\n_{_\n\n_\"meta\": {_\n\n_\"warnings\": \\[\\],_\n\n_\"page\": {_\n\n_\"current\": 11,_\n\n_\"total\\_pages\": 0,_\n\n_\"total\\_results\": 0,_\n\n_\"size\": 1000_\n\n_},_\n\n_\"request\\_id\": \"39684c716fe14725a70406a1a71789e4\"_\n\n_},_\n\n_\"results\": \\[\\] <--- 0 results here_\n\n_}_\n\nSuddenly it indicates that there are no docs found at all\u2026\n\nThe documentation says that search-request should support up to 1000 in page size and up to 500 pages, but it seems like it supports max 10 pages when page-size is 1000. Or is there some setting I need to change to support more result-pages?\n\nOr is there some other way I can request ids of all docs in the index?\n\nAny help would be appreciated\n\n2.1k\nviews\n\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)3](https://discuss.elastic.co/u/frankjoh2 \"frankjoh2\")\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)2](https://discuss.elastic.co/u/goodroot \"goodroot\")\n\n## post by frankjoh2 on Dec 7, 2018\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)](https://discuss.elastic.co/u/frankjoh2)\n\n[frankjoh2](https://discuss.elastic.co/u/frankjoh2)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/2 \"Post date\")\n\nI found out now that there is a list-method (/documents/list) this is specifically meant to get all docs in an index. But this will return all data (not just the id-field) and has a max pagesize of 100 docs.\n\nThis means I'll have to spend at least 10 times as many api-requests to get this done and that might push me above the monthly limit and resulting in extra licensing-costs.\n\nSo I'd still be interested to know of any alternatives if they exist.\n\n## post by goodroot on Dec 7, 2018\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)](https://discuss.elastic.co/u/goodroot)\n\n[goodroot](https://discuss.elastic.co/u/goodroot)[Kellen Evan](https://discuss.elastic.co/u/goodroot)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/3 \"Post date\")\n\nFrank!\n\nYou are correct. The most effective method of returning documents would be to iterate over the `documents/list` endpoint, like so:\n\n```rust\n\ncurl -X GET 'https://host-xxxxxx.api.swiftype.com/api/as/v1/engines/example-engine/documents/list' \\\n-H 'Content-Type: application/json' \\\n-H 'Authorization: Bearer private-xxxxxxxxxxxxxxxxxxxx' \\\n-d '{\n  \"page\": {\n    \"current\": 1,\n    \"size\": 100\n  }\n}'\n```\n\nAs you pointed out, this will return full documents, not just the `id`. You will need to parse out the `id` when assembling your list.\n\nThanks for posting, I wish you an excellent end to your week.\n\n## post by frankjoh2 on Dec 10, 2018\n\n[![](https://avatars.discourse-cdn.com/v4/letter/f/85f322/48.png)](https://discuss.elastic.co/u/frankjoh2)\n\n[frankjoh2](https://discuss.elastic.co/u/frankjoh2)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/4 \"Post date\")\n\nHi goodroot and thanks for the reply. I tried to use the documents/list endpoint now and sadly that doesn't work either. When the \"current\"-attribute in the request gets higher than 100 the response always returns the 100th page. In other words, it's not possible to get more than the first 10.000 documents (100 pages with 100 documents each) with this endpoint too.\n\n## post by goodroot on Dec 10, 2018\n\n[![](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/goodroot/48/38389_2.png)](https://discuss.elastic.co/u/goodroot)\n\n[goodroot](https://discuss.elastic.co/u/goodroot)[Kellen Evan](https://discuss.elastic.co/u/goodroot)\n\n[Dec 2018](https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918/5 \"Post date\")\n\nFrank --\n\nI poked around to try and find you a better answer, but that is correct.\n\nThe limit of both `current` and `size` is `100`, and the endpoint cannot be used to return more than `10,000` documents.\n\nI understand this creates a gap. The documents, and their ids, are available by querying or through the documents dashboard view, but that isn't helpful to one looking to generate a comprehensive list. ![:confused:](https://emoji.discourse-cdn.com/twitter/confused.png?v=6)\n\nThe limit may rise in the future, but as of now it is kept restrictive. If this is a major blocker in your use-case, please email [support@swiftype.com](mailto:support@swiftype.com), referencing this ticket so that we can learn more.\n\nEnjoy the week,\n\nKellen\n\n28 days later\n\n\n## Closed on Jan 7, 2019\n\n[![](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png)](https://discuss.elastic.co/u/system)\n\nClosed on Jan 7, 2019\n\nThis topic was automatically closed 28 days after the last reply. New replies are no longer allowed.\n\nReply\n\n### Related topics\n\n| Topic | Replies | Views | Activity |\n| --- | --- | --- | --- |\n| [Getting next 10k documents with AppSearch.list\\_documents()](https://discuss.elastic.co/t/getting-next-10k-documents-with-appsearch-list-documents/345216)<br>[Elasticsearch](https://discuss.elastic.co/c/elastic-stack/elasticsearch/6) | [5](https://discuss.elastic.co/t/getting-next-10k-documents-with-appsearch-list-documents/345216/1) | 254 | [Oct 2023](https://discuss.elastic.co/t/getting-next-10k-documents-with-appsearch-list-documents/345216/6) |\n| [Results per Page limit](https://discuss.elastic.co/t/results-per-page-limit/217643)<br>[Elastic Search](https://discuss.elastic.co/c/search/84) <br>- [elastic-app-search](https://discuss.elastic.co/tag/elastic-app-search/118) | [5](https://discuss.elastic.co/t/results-per-page-limit/217643/1) | 573 | [Feb 2020](https://discuss.elastic.co/t/results-per-page-limit/217643/6) |\n| [Get all ids with Python](https://discuss.elastic.co/t/get-all-ids-with-python/344689)<br>[Elasticsearch](https://discuss.elastic.co/c/elastic-stack/elasticsearch/6) | [1](https://discuss.elastic.co/t/get-all-ids-with-python/344689/1) | 394 | [Oct 2023](https://discuss.elastic.co/t/get-all-ids-with-python/344689/2) |\n| [How to get results over 10K in App search](https://discuss.elastic.co/t/how-to-get-results-over-10k-in-app-search/351310)<br>[Elastic Search](https://discuss.elastic.co/c/search/84) <br>- [elastic-app-search](https://discuss.elastic.co/tag/elastic-app-search/118) | [10](https://discuss.elastic.co/t/how-to-get-results-over-10k-in-app-search/351310/1) | 13.3k | [Jan 2024](https://discuss.elastic.co/t/how-to-get-results-over-10k-in-app-search/351310/11) |\n| [I need to fetch all document but this query only return 10 documents. I\u2019m checking through postman](https://discuss.elastic.co/t/i-need-to-fetch-all-document-but-this-query-only-return-10-documents-im-checking-through-postman/229574)<br>[Elasticsearch](https://discuss.elastic.co/c/elastic-stack/elasticsearch/6) | [1](https://discuss.elastic.co/t/i-need-to-fetch-all-document-but-this-query-only-return-10-documents-im-checking-through-postman/229574/1) | 1.1k | [Apr 2020](https://discuss.elastic.co/t/i-need-to-fetch-all-document-but-this-query-only-return-10-documents-im-checking-through-postman/229574/2) |\n\nTopic list, column headers with buttons are sortable.\n\n\u00a9 2020\\. All Rights Reserved - Elasticsearch\n\n- Elasticsearch is a trademark of Elasticsearch BV, registered in the U.S.\nand in other countries\n\n- [Trademarks](https://www.elastic.co/legal/trademarks)\n- [Terms](https://www.elastic.co/legal/terms-of-use)\n- [Privacy](https://www.elastic.co/legal/privacy-policy)\n- [Brand](https://www.elastic.co/brand)\n- [Code of Conduct](https://www.elastic.co/community/codeofconduct)\n\nApache, Apache Lucene, Apache Hadoop, Hadoop, HDFS and the yellow elephant\nlogo are trademarks of the\n[Apache Software Foundation](http://www.apache.org/)\nin the United States and/or other\u00a0countries.",
          "metadata": {
            "ogImage": "https://us1.discourse-cdn.com/elastic/original/3X/5/4/5461df8fd2fe783981b0180332821184b729980e.png",
            "article:published_time": "2018-12-07T11:22:57+00:00",
            "og:ignore_canonical": "true",
            "theme-color": [
              "#ffffff",
              "#141619",
              "#ffffff"
            ],
            "ogTitle": "Fetching all docs in an app search index - Elastic Search - Discuss the Elastic Stack",
            "title": "Fetching all docs in an app search index - Elastic Search - Discuss the Elastic Stack",
            "publishedTime": "2018-12-07T11:22:57+00:00",
            "og:article:section": "Elastic Search",
            "og:description": "Hi,  I'm trying to get the ids of all documents in my app search index. I tried to simply iteration of searches with an empty query and incrementally increasing the page-number. This seemed to work fine for the first 10 requests.  This is what my requests looks like:  {  \"query\": \"\",  \"result_fields\": {  \"id\": { \"raw\": {} }  },  \"page\": {\"size\": 1000,\"current\": [PAGENUMER] }  }  Where [PAGENUMBER] is 1 for the first request, 2 for second and so on\u2026  This is the result of the 10th request:  {  \"m...",
            "discourse_current_homepage": "categories",
            "ogDescription": "Hi,  I'm trying to get the ids of all documents in my app search index. I tried to simply iteration of searches with an empty query and incrementally increasing the page-number. This seemed to work fine for the first 10 requests.  This is what my requests looks like:  {  \"query\": \"\",  \"result_fields\": {  \"id\": { \"raw\": {} }  },  \"page\": {\"size\": 1000,\"current\": [PAGENUMER] }  }  Where [PAGENUMBER] is 1 for the first request, 2 for second and so on\u2026  This is the result of the 10th request:  {  \"m...",
            "google-site-verification": [
              "DomjkVg_vXHmIOm6fe3rZEhQ5tCXZl9fb-H-SNkCKaY",
              "ifAAE_mkvVVDx_RVqNrCqZMXKqZKxdt_XZIYL33pjtE"
            ],
            "twitter:url": "https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918",
            "ogUrl": "https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918",
            "color-scheme": "light dark",
            "og:site_name": "Discuss the Elastic Stack",
            "ogSiteName": "Discuss the Elastic Stack",
            "twitter:image": "https://us1.discourse-cdn.com/elastic/original/3X/5/4/5461df8fd2fe783981b0180332821184b729980e.png",
            "twitter:description": "Hi,  I'm trying to get the ids of all documents in my app search index. I tried to simply iteration of searches with an empty query and incrementally increasing the page-number. This seemed to work fine for the first 10 requests.  This is what my requests looks like:  {  \"query\": \"\",  \"result_fields\": {  \"id\": { \"raw\": {} }  },  \"page\": {\"size\": 1000,\"current\": [PAGENUMER] }  }  Where [PAGENUMBER] is 1 for the first request, 2 for second and so on\u2026  This is the result of the 10th request:  {  \"m...",
            "og:type": "website",
            "og:article:tag": "elastic-app-search",
            "msapplication-TileColor": "#ffffff",
            "og:url": "https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918",
            "og:article:section:color": "0088CC",
            "language": "en",
            "generator": "Discourse 2026.9.0-latest - https://github.com/discourse/discourse version dea480002bf1fa9315c4022892faf16a9902c8e5",
            "discourse-track-view-session-id": "QH1glfVJSFKBO7bure6p1C2SaqEE8roC",
            "discourse-beacon-pageview-enabled": "true",
            "twitter:card": "summary",
            "discourse-engagement-tracking-enabled": "true",
            "twitter:title": "Fetching all docs in an app search index - Elastic Search - Discuss the Elastic Stack",
            "msapplication-TileImage": "https://www.elastic.co/ms-icon-144x144.png",
            "description": "Hi, \nI'm trying to get the ids of all documents in my app search index. I tried to simply iteration of searches with an empty query and incrementally increasing the page-number. This seemed to work fine for the first 10 \u2026",
            "fragment": "!",
            "discourse_theme_id": "25",
            "og:image": "https://us1.discourse-cdn.com/elastic/original/3X/5/4/5461df8fd2fe783981b0180332821184b729980e.png",
            "discourse/config/environment": "%7B%22modulePrefix%22%3A%22discourse%22%2C%22environment%22%3A%22production%22%2C%22rootURL%22%3A%22%22%2C%22locationType%22%3A%22history%22%2C%22EmberENV%22%3A%7B%22FEATURES%22%3A%7B%7D%2C%22EXTEND_PROTOTYPES%22%3Afalse%7D%2C%22APP%22%3A%7B%22name%22%3A%22discourse%22%2C%22version%22%3A%222026.9.0-latest%20dea480002bf1fa9315c4022892faf16a9902c8e5%22%7D%7D",
            "og:title": "Fetching all docs in an app search index - Elastic Search - Discuss the Elastic Stack",
            "viewport": "width=device-width, initial-scale=1.0, minimum-scale=1.0, viewport-fit=cover, interactive-widget=resizes-content",
            "favicon": "https://us1.discourse-cdn.com/elastic/optimized/3X/3/7/37d58c6890e29b4d96eae5c67db6cf3db36b0181_2_32x32.png",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-a37456b0e2e4",
            "sourceURL": "https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918",
            "url": "https://discuss.elastic.co/t/fetching-all-docs-in-an-app-search-index/159918",
            "statusCode": 200,
            "contentType": "text/html; charset=utf-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "56707a84-15e1-4cc8-9d44-e6078777be0a",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://medium.com/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603",
          "title": "Turning the Web Into a Real-Time Database with OloStep - Medium",
          "description": "OloStep turns the open web into a real-time knowledge source structured, filterable, and directly useful for agents, workflows, and apps. Let's ...",
          "position": 6,
          "markdown": "[Sitemap](https://medium.com/sitemap/sitemap.xml)\n\n[Open in app](https://play.google.com/store/apps/details?id=com.medium.reader&referrer=utm_source%3DmobileNavBar&source=---top_nav_layout_nav-----------------------------------------)\n\nSign up\n\n[Sign in](https://medium.com/m/signin?operation=login&redirect=https%3A%2F%2Fmedium.com%2F%40davidfagb%2Fturning-the-web-into-a-real-time-database-with-olostep-f175e956a603&source=post_page---top_nav_layout_nav-----------------------global_nav------------------)\n\n[Medium Logo](https://medium.com/?source=---top_nav_layout_nav-----------------------------------------)\n\nGet app\n\n[Write](https://medium.com/m/signin?operation=register&redirect=https%3A%2F%2Fmedium.com%2Fnew-story&source=---top_nav_layout_nav-----------------------new_post_topnav------------------)\n\n[Search](https://medium.com/search?source=---top_nav_layout_nav-----------------------------------------)\n\nSign up\n\n[Sign in](https://medium.com/m/signin?operation=login&redirect=https%3A%2F%2Fmedium.com%2F%40davidfagb%2Fturning-the-web-into-a-real-time-database-with-olostep-f175e956a603&source=post_page---top_nav_layout_nav-----------------------global_nav------------------)\n\n![Unknown user](https://miro.medium.com/v2/resize:fill:32:32/1*dmbNkD5D-u45r44go_cf0g.png)\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:40:40/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---post_author_sidebar--f175e956a603-----------------d1058d8a6759----------------------)\n\n## David Fagbuyiro\n\nTechnical writer\n\nFollow writer\n\n[Web Scraping](https://medium.com/tag/web-scraping?source=post_page---header_tags--f175e956a603---------------------------------------)\n\n[Database](https://medium.com/tag/database?source=post_page---header_tags--f175e956a603---------------------------------------)\n\n[Web](https://medium.com/tag/web?source=post_page---header_tags--f175e956a603---------------------------------------)\n\n# Turning the Web Into a Real-Time Database with OloStep\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:32:32/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---byline--f175e956a603---------------------------------------)\n\n[David Fagbuyiro](https://medium.com/@davidfagb?source=post_page---byline--f175e956a603---------------------------------------)\n\nFollow\n\n4 min read\n\n\u00b7\n\nOct 8, 2025\n\n3\n\n[Listen](https://medium.com/m/signin?actionUrl=https%3A%2F%2Fmedium.com%2Fplans%3Fdimension%3Dpost_audio_button%26postId%3Df175e956a603&operation=register&redirect=https%3A%2F%2Fmedium.com%2F%40davidfagb%2Fturning-the-web-into-a-real-time-database-with-olostep-f175e956a603&source=---header_actions--f175e956a603---------------------post_audio_button------------------)\n\nShare\n\nYou\u2019re building an AI assistant for a startup CEO who needs to monitor competitors, track pricing updates, new hires, and product launches. Straightforward stuff. So, you start hunting for data scraping solutions that are too brittle to break after a minor site change. Most sites do not offer APIs. Google? Buried in ads, captchas, and unstructured results. You hit a wall.\n\nThen it hits you that the data is all out there, it\u2019s just trapped inside messy webpages. That\u2019s precisely what OloStep is built to fix. Instead of scraping, crawling, or building static indexes, OloStep flips the model. It treats the open web itself as a live, structured, and queryable database in real time.\n\nPress enter or click to view image in full size\n\n![Illustration of a chaotic web full of HTML pages turning into a clean, structured database view](https://miro.medium.com/v2/resize:fit:700/1*kCHsAMclE1H0wYXNZjqlIA.png)\n\n### The Web Wasn\u2019t Meant to Be Queried\n\nThe internet is one of the most extensive and valuable datasets in history. However, it was not designed for structure, but it was created for publishing.\n\nThere are billions of pages, including blogs, product listings, pricing tables, and job boards, which are constantly being updated. Yet, querying this information in a traditional manner is nearly impossible. You can crawl, index, and scrape the web, but these methods are often brittle, slow, and quickly become outdated.\n\nWhat if, instead, you could treat the open web as if it were your live backend?\n\n## What is Olostep?\n\nOlostep is a powerful web scraping API that efficiently provides e data from any website. It meets the essential demands for fast, reliable, and cost-effective data acquisition, AI development and large-scale data aggregation. rtups, and established companies enhance their data-driven applications and automate workflows.\n\n### The Core Idea: Query the Open Web Like a Database\n\nThink SQL, but the tables are web pages. Think APIs, but the endpoints are search queries. OloStep turns the open web into a real-time knowledge source structured, filterable, and directly useful for agents, workflows, and apps.\n\nLet\u2019s say you\u2019re building an AI agent that finds pricing trends for B2B SaaS tools. Traditionally, you\u2019d:\n\n- Crawl pages\n- Parse them manually\n- Store them\n- Run queries later\n\nBut with OloStep:\n\n```\nSELECT price, plan_name, vendor\nFROM web\nWHERE category = \"CRM software\" AND updated_within = \"7 days\"\n```\n\nIt\u2019s not a stretch. This is the direction we\u2019re heading to obtain structured results directly from the public internet, without scraping or even relying on pipelines, just answers.\n\nPress enter or click to view image in full size\n\n![Side-by-side comparison\u200a\u2014\u200aLeft: messy HTML page, Right: clean JSON-like table extracted live].](https://miro.medium.com/v2/resize:fit:700/1*NL3okKxPHQ4mnYe52GbmvQ.png)\n\nSide-by-side comparison \u2014 Left: messy HTML page, Right: clean JSON-like table extracted live\\].\n\n### What Makes OloStep Different\n\nMost tools that extract data from the web fall into one of two buckets:\n\n1. **Scrapers**: Fast but fragile. Break often, need constant maintenance.\n2. **Knowledge graphs:** Structured but stale. Require time-consuming ingestion pipelines.\n\nOloStep is different because it utilizes AI-native tools to extract structured data pages in real **-** time. The model understands both layout and meaning. It can:\n\n- Recognize product specs on landing pages\n- Pull out job listings across multiple sites\n- Extract facts from articles or documentation\n\nIt\u2019s a web-native query engine. No scraping rules. No brittle XPath selectors. If you\u2019re building anything that depends on real-world knowledge, you don\u2019t want to be stuck waiting for someone else to scrape, ingest, and update it for you.\n\nPress enter or click to view image in full size\n\n![Flowchart showing an AI agent querying OloStep -> real-time web pages -> structured results](https://miro.medium.com/v2/resize:fit:700/1*jl9B2mmsvoodZF5HbGzcoA.png)\n\nFlowchart showing an AI agent querying OloStep -> real-time web pages -> structured results\n\n### Why Now? Because the Stack Just Clicked\n\nThis wasn\u2019t possible five years ago.\n\n- **LLMs** were too weak.\n- **Scrapers** were too brittle.\n- **Knowledge graphs** were too slow.\n\nHowever, now LLMs can comprehend, reason over tables, and infer meaning across web pages in real-time.\n\n## Get David Fagbuyiro\u2019s stories in\u00a0your\u00a0inbox\n\nJoin Medium for free to get updates from\u00a0this\u00a0writer.\n\nSubscribe\n\nSubscribe\n\nRemember me for faster sign in\n\nOloStep rides that wave, turning raw web content into structured, semantically rich data at the moment you need it. It\u2019s not just search. It\u2019s understanding **.**\n\n## Use Cases Already Emerging\n\nBelow are the benefits of implementing OloStep:\n\n- AI Research Agents: agents that answer questions by citing live data from across the web, not just summaries from cached pages.\n- Competitive Intelligence: pull pricing changes, product launches, or staffing shifts from public pages and career portals.\n- Custom Tools & Dashboards: develop internal and public-facing information, such as public tenders, vendor updates, or compliance announcements.\n- Programmatic Search: replace brittle Google scraping with structured queries and filters. No hacks. No captchas. Just answers.\n\n### The Future: Don\u2019t Store the Web. Ask It.\n\nOloStep flips the traditional methods so instead of trying to tame the chaos of the internet into a static database, it embraces the mess and makes it readable, understandable, and usable.\n\nThink of it like this:\n\n- Search gives you links.\n- Scraping gives you fragments.\n- OloStep provides you with facts, and it does so in real-time.\n\nSo instead of building brittle scraping tools or waiting for someone to publish a dataset, you can now ask the web a question and get structured, filtered, useful data in return.\n\nThe web is no longer just something you read; it\u2019s something you query.\n\n### Try It Yourself\n\nIf you\u2019re building agents, dashboards, or any product that depends on fresh, structured information from the real world, OloStep gives you the edge.\n\nCheck it out at [https://olostep.com](https://olostep.com/)\n\nStart asking the web real questions and get real answers.\n\nWith OloStep, you don\u2019t have to.\n\n[Web Scraping](https://medium.com/tag/web-scraping?source=post_page---footer_tags--f175e956a603---------------------------------------)\n\n[Database](https://medium.com/tag/database?source=post_page---footer_tags--f175e956a603---------------------------------------)\n\n[Web](https://medium.com/tag/web?source=post_page---footer_tags--f175e956a603---------------------------------------)\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:48:48/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\n[![David Fagbuyiro](https://miro.medium.com/v2/resize:fill:64:64/1*tQTjIfY643Gzl6ovvctb4A.jpeg)](https://medium.com/@davidfagb?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\nFollow\n\n[**Written by David Fagbuyiro**](https://medium.com/@davidfagb?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\n[54 followers](https://medium.com/@davidfagb/followers?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\n\u00b7 [6 following](https://medium.com/@davidfagb/following?source=post_page---post_author_info--f175e956a603---------------------------------------)\n\nTechnical writer\n\nFollow\n\n[Help](https://help.medium.com/hc/en-us?source=post_page-----f175e956a603---------------------------------------)\n\n[Status](https://status.medium.com/?source=post_page-----f175e956a603---------------------------------------)\n\n[About](https://medium.com/about?autoplay=1&source=post_page-----f175e956a603---------------------------------------)\n\n[Careers](https://medium.com/jobs-at-medium/work-at-medium-959d1a85284e?source=post_page-----f175e956a603---------------------------------------)\n\n[Press](mailto:pressinquiries@medium.com)\n\n[Blog](https://blog.medium.com/?source=post_page-----f175e956a603---------------------------------------)\n\n[Store](https://medium.com/store)\n\n[Privacy](https://policy.medium.com/medium-privacy-policy-f03bf92035c9?source=post_page-----f175e956a603---------------------------------------)\n\n[Rules](https://policy.medium.com/medium-rules-30e5502c4eb4?source=post_page-----f175e956a603---------------------------------------)\n\n[Terms](https://policy.medium.com/medium-terms-of-service-9db0094a1e0f?source=post_page-----f175e956a603---------------------------------------)\n\n[Text to speech](https://speechify.com/medium?source=post_page-----f175e956a603---------------------------------------)",
          "metadata": {
            "robots": "index,follow,max-image-preview:large",
            "og:description": "You\u2019re building an AI assistant for a startup CEO who needs to monitor competitors, track pricing updates, new hires, and product launches\u2026",
            "twitter:app:id:iphone": "828256236",
            "ogDescription": "You\u2019re building an AI assistant for a startup CEO who needs to monitor competitors, track pricing updates, new hires, and product launches\u2026",
            "title": "Turning the Web Into a Real-Time Database with OloStep | by David Fagbuyiro | Medium",
            "al:ios:url": "medium://p/f175e956a603",
            "publishedTime": "2025-10-08T13:29:32.242Z",
            "al:android:package": "com.medium.reader",
            "al:web:url": "https://medium.com/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603",
            "og:image": "https://miro.medium.com/v2/resize:fit:1200/1*kCHsAMclE1H0wYXNZjqlIA.png",
            "twitter:app:url:iphone": "medium://p/f175e956a603",
            "twitter:description": "You\u2019re building an AI assistant for a startup CEO who needs to monitor competitors, track pricing updates, new hires, and product launches\u2026",
            "al:ios:app_store_id": "828256236",
            "description": "Turning the Web Into a Real-Time Database with OloStep You\u2019re building an AI assistant for a startup CEO who needs to monitor competitors, track pricing updates, new hires, and product launches \u2026",
            "ogUrl": "https://medium.com/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603",
            "author": "David Fagbuyiro",
            "referrer": "unsafe-url",
            "twitter:image:alt": "Illustration of a chaotic web full of HTML pages turning into a clean, structured database view",
            "twitter:creator": "@Davidfagb",
            "twitter:title": "Turning the Web Into a Real-Time Database with OloStep",
            "twitter:data1": "4 min read",
            "al:ios:app_name": "Medium",
            "apple-itunes-app": "app-id=828256236, app-argument=/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603, affiliate-data=pt=698524&ct=smart_app_banner&mt=8",
            "og:url": "https://medium.com/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603",
            "language": "en",
            "theme-color": "#000000",
            "twitter:site": "@Medium",
            "ogSiteName": "Medium",
            "og:image:alt": "Illustration of a chaotic web full of HTML pages turning into a clean, structured database view",
            "article:author": "https://medium.com/@davidfagb",
            "twitter:label1": "Reading time",
            "fb:app_id": "542599432471018",
            "og:site_name": "Medium",
            "og:title": "Turning the Web Into a Real-Time Database with OloStep",
            "twitter:image:src": "https://miro.medium.com/v2/resize:fit:1200/1*kCHsAMclE1H0wYXNZjqlIA.png",
            "og:type": "article",
            "al:android:url": "medium://p/f175e956a603",
            "viewport": [
              "width=device-width, initial-scale=1",
              "width=device-width,minimum-scale=1,initial-scale=1,maximum-scale=1"
            ],
            "ogImage": "https://miro.medium.com/v2/resize:fit:1200/1*kCHsAMclE1H0wYXNZjqlIA.png",
            "twitter:app:name:iphone": "Medium",
            "twitter:card": "summary_large_image",
            "ogTitle": "Turning the Web Into a Real-Time Database with OloStep",
            "article:published_time": "2025-10-08T13:29:32.242Z",
            "al:android:app_name": [
              "Medium",
              "Medium"
            ],
            "favicon": "https://miro.medium.com/v2/5d8de952517e8160e40ef9841c781cdc14a5db313057fa3c3de41c6f5b494b19",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-a4bcc9a58d99",
            "sourceURL": "https://medium.com/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603",
            "url": "https://medium.com/@davidfagb/turning-the-web-into-a-real-time-database-with-olostep-f175e956a603",
            "statusCode": 200,
            "contentType": "text/html; charset=utf-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "0e8bc2fe-1cce-45bc-9c01-a563302040d7",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://docs.olostep.com/get-started/welcome",
          "title": "Welcome to Olostep - Olostep Docs",
          "description": "## [\u200b](https://docs.olostep.com/get-started/welcome#scrape)  Scrape\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\nresult = client.scrapes.create(\n    url_to_scrape=\"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n    formats=[\"markdown\"],\n)\nprint(result.markdown_content)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst result = await client.scrapes.create({\n  url: 'https://en.wikipedia.org/wiki/Alexander_the_Great',\n  formats: ['markdown'],\n})\nconsole.log(result.markdown_content)\n```\n\n## [\u200b](https://docs.olostep.com/get-started/welcome#crawl)  Crawl\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\ncrawl = client.crawls.create(\n    start_url=\"https://docs.olostep.com\",\n    max_pages=50,\n)\nfor page in crawl.pages():\n    print(page.url)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst crawl = await client.crawls.create({\n  url: 'https://docs.olostep.com',\n  maxPages: 50,\n})\nfor await (const page of crawl.pages()) {\n  console.log(page.url)\n}\n```\n\n```\nolostep crawl \"https://docs.olostep.com\" --max-pages 50\n```\n\n## [\u200b](https://docs.olostep.com/get-started/welcome#answer)  Answer\n```\nolostep answer \"What does Olostep do?\"\n```",
          "position": 7,
          "markdown": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/get-started/welcome#content-area)\n\nOlostep is infrastructure for the Web\u2019s second user \u2014 giving AI agents a way to search the web, extract structured data in real time, and build custom research agents.![Main dashboard interface](https://mintcdn.com/olostep-58/ndWPtiE100GTiImQ/images/explainer_v5.png?fit=max&auto=format&n=ndWPtiE100GTiImQ&q=85&s=ae92c6f02390b885a46f830090b0de34)\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#introduction)  Introduction\n\n- The **Olostep API** is the best **web search**, **scraping** and **crawling** API for AI used by some of the leading startups and scaleups in the world.\n- The **Olostep Agent** allows to automate **research workflows** in a no code way with just a prompt in natural language.\n\n**For AI agents:** fetch [docs.olostep.com/llms.txt](https://docs.olostep.com/llms.txt) for a complete index of this documentation before exploring further or [get started here](https://www.olostep.com/agent-onboarding/SKILL.md).\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#use-olostep-from-your-terminal-and-ai-agents)  Use Olostep from your terminal and AI agents\n\nBeyond the API, Olostep ships a CLI, an MCP server, and drop-in skills so any tool \u2014 Claude Code, Cursor, Windsurf, and more \u2014 can use the web natively.\n\n[**CLI \u2192** \\\\\n\\\\\n`npm i -g olostep-cli` \u2014 scrape, map, crawl, answer, and batch the web from your terminal. JSON output for scripts, CI, and agents.](https://docs.olostep.com/sdks/cli)\n\n[**MCP Server \u2192** \\\\\n\\\\\nGive any MCP client (Claude, Cursor, VS Code) live web tools. Hosted endpoint \u2014 no install.](https://docs.olostep.com/integrations/mcp-server)\n\n[**Skills \u2192** \\\\\n\\\\\nDrop-in skills that teach AI coding agents how and when to use Olostep. Install with `olostep add skills`.](https://docs.olostep.com/features/skills)\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#what-can-olostep-do)  What can Olostep do?\n\n[**Scrape** \\\\\n\\\\\nPull any URL as clean Markdown, HTML, screenshots, or structured JSON.](https://docs.olostep.com/get-started/welcome#scrape)\n\n[**Crawl** \\\\\n\\\\\nRecursively gather every page on a site, with filters and search.](https://docs.olostep.com/get-started/welcome#crawl)\n\n[**Answer** \\\\\n\\\\\nGet AI-synthesised answers from live web sources, with citations.](https://docs.olostep.com/get-started/welcome#answer)\n\n### [\u200b](https://docs.olostep.com/get-started/welcome\\#why-olostep)  Why Olostep?\n\n- **Built for AI**: Clean Markdown, structured JSON, citations \u2014 output your agents and apps consume directly.\n- **Reliable at scale**: Industry-leading success rate; handles JavaScript, anti-bot, and proxies under the hood.\n- **Fast**: Sub-second single scrape; up to 10,000 URLs in a single batch in 5\u20137 minutes.\n- **Cost-effective**: Significantly cheaper than alternatives at production scale.\n- **CLI + MCP + Skills**: Use Olostep from your terminal, scripts, or any MCP-aware agent \u2014 agent skills included.\n\n* * *\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#scrape)  Scrape\n\nPull any URL as clean Markdown. See the [Scrape feature docs](https://docs.olostep.com/features/scrapes) for all options.\n\nPython\n\nNode\n\ncURL\n\nCLI\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\nresult = client.scrapes.create(\n    url_to_scrape=\"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n    formats=[\"markdown\"],\n)\nprint(result.markdown_content)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst result = await client.scrapes.create({\n  url: 'https://en.wikipedia.org/wiki/Alexander_the_Great',\n  formats: ['markdown'],\n})\nconsole.log(result.markdown_content)\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/scrapes\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"url_to_scrape\": \"https://en.wikipedia.org/wiki/Alexander_the_Great\",\n    \"formats\": [\"markdown\"]\n  }'\n```\n\n```\nolostep scrape \"https://en.wikipedia.org/wiki/Alexander_the_Great\"\n```\n\nResponse\n\n```\n{\n  \"id\": \"scrape_6h89o8u1kt\",\n  \"object\": \"scrape\",\n  \"result\": {\n    \"markdown_content\": \"## Alexander the Great...\",\n    \"markdown_hosted_url\": \"https://olostep-storage.s3.us-east-1.amazonaws.com/markDown_6h89o8u1kt.txt\",\n    \"page_metadata\": { \"status_code\": 200, \"title\": \"Alexander the Great - Wikipedia\" }\n  }\n}\n```\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#crawl)  Crawl\n\nRecursively gather every page on a site, with include/exclude filters and an optional `search_query` to focus the crawl. See [Crawl feature docs](https://docs.olostep.com/features/crawls).\n\nPython\n\nNode\n\ncURL\n\nCLI\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\ncrawl = client.crawls.create(\n    start_url=\"https://docs.olostep.com\",\n    max_pages=50,\n)\nfor page in crawl.pages():\n    print(page.url)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst crawl = await client.crawls.create({\n  url: 'https://docs.olostep.com',\n  maxPages: 50,\n})\nfor await (const page of crawl.pages()) {\n  console.log(page.url)\n}\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/crawls\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"start_url\": \"https://docs.olostep.com\",\n    \"max_pages\": 50\n  }'\n```\n\n```\nolostep crawl \"https://docs.olostep.com\" --max-pages 50\n```\n\nResponse\n\n```\n{\n  \"id\": \"crawl_abc123\",\n  \"object\": \"crawl\",\n  \"status\": \"completed\",\n  \"pages_count\": 47,\n  \"pages\": [\\\n    { \"url\": \"https://docs.olostep.com/get-started/welcome\", \"retrieve_id\": \"...\" },\\\n    { \"url\": \"https://docs.olostep.com/features/scrapes\", \"retrieve_id\": \"...\" }\\\n  ]\n}\n```\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#answer)  Answer\n\nAsk a question and get an AI-synthesised answer from live web sources, with citations. Pass a JSON schema to shape the output. See [Answers docs](https://docs.olostep.com/features/answers).\n\nPython\n\nNode\n\ncURL\n\nCLI\n\n```\nfrom olostep import Olostep\n\nclient = Olostep(api_key=\"YOUR_REAL_KEY\")\nanswer = client.answers.create(task=\"What does Olostep do?\")\nprint(answer.result)\nprint(answer.sources)\n```\n\n```\nimport Olostep from 'olostep'\n\nconst client = new Olostep({ apiKey: 'YOUR_REAL_KEY' })\nconst answer = await client.answers.create({ task: 'What does Olostep do?' })\nconsole.log(answer.result)\nconsole.log(answer.sources)\n```\n\n```\ncurl -s -X POST \"https://api.olostep.com/v1/answers\" \\\n  -H \"Authorization: Bearer $OLOSTEP_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"task\": \"What does Olostep do?\"}'\n```\n\n```\nolostep answer \"What does Olostep do?\"\n```\n\nResponse\n\n```\n{\n  \"id\": \"answer_abc123\",\n  \"object\": \"answer\",\n  \"task\": \"What does Olostep do?\",\n  \"result\": {\n    \"json_content\": \"{\\\"result\\\":\\\"Olostep is an API that lets AI agents search, scrape, and structure web data.\\\"}\",\n    \"sources\": [\\\n      \"https://docs.olostep.com/get-started/welcome\",\\\n      \"https://www.olostep.com/\"\\\n    ]\n  }\n}\n```\n\n* * *\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#more-capabilities)  More capabilities\n\n[**Batch** \\\\\n\\\\\nScrape up to 10,000 URLs in parallel; results back in 5\u20137 minutes.](https://docs.olostep.com/features/batches)\n\n[**Map** \\\\\n\\\\\nDiscover every URL on a site with include/exclude patterns.](https://docs.olostep.com/features/maps)\n\n[**Search** \\\\\n\\\\\nLive web search with structured links and optional inline scraping.](https://docs.olostep.com/features/search)\n\n[**Parsers** \\\\\n\\\\\nSelf-healing extractors that turn pages into typed JSON at scale.](https://docs.olostep.com/features/structured-content/parsers)\n\n[**Schedules** \\\\\n\\\\\nRun scrapes, crawls, and answers on a recurring schedule.](https://docs.olostep.com/features/schedules)\n\n[**Files** \\\\\n\\\\\nUpload files for batches or to connect your knowledge base.](https://docs.olostep.com/features/files)\n\n* * *\n\n## [\u200b](https://docs.olostep.com/get-started/welcome\\#resources)  Resources\n\n[**Explore Features** \\\\\n\\\\\nCheck out all supported features for your scraping and AI search needs.](https://docs.olostep.com/features/scrapes)\n\n[**API Reference** \\\\\n\\\\\nStart using the API and test out the various params.](https://docs.olostep.com/api-reference/scrapes/create)\n\n[**Integrations** \\\\\n\\\\\nUse Olostep in n8n, make, relay, zapier, etc](https://docs.olostep.com/integrations/n8n)\n\n[**Examples** \\\\\n\\\\\nBrowse ready-to-use examples to get started quickly.](https://docs.olostep.com/examples/)\n\nWas this page helpful?\n\nYesNo",
          "metadata": {
            "msapplication-TileColor": "#9563FF",
            "description": "Olostep: Infrastructure for the Web's second user. The best search, scraping and crawling API for AI.",
            "twitter:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DGet%2BStarted%26title%3DWelcome%2Bto%2BOlostep%26description%3DOlostep%253A%2BInfrastructure%2Bfor%2Bthe%2BWeb%2527s%2Bsecond%2Buser.%2BThe%2Bbest%2Bsearch%252C%2Bscraping%2Band%2Bcrawling%2BAPI%2Bfor%2BAI.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "og:title": "Welcome to Olostep - Olostep Docs",
            "ogTitle": "Welcome to Olostep - Olostep Docs",
            "application-name": "Olostep Docs",
            "og:description": "Olostep: Infrastructure for the Web's second user. The best search, scraping and crawling API for AI.",
            "og:image:height": "630",
            "og:type": "website",
            "twitter:description": "Olostep: Infrastructure for the Web's second user. The best search, scraping and crawling API for AI.",
            "ogUrl": "https://docs.olostep.com/get-started/welcome",
            "apple-mobile-web-app-title": "Olostep Docs",
            "googlebot": "index, follow",
            "language": "en",
            "robots": "index, follow",
            "title": "Welcome to Olostep - Olostep Docs",
            "viewport": "width=device-width, initial-scale=1, viewport-fit=cover",
            "og:url": "https://docs.olostep.com/get-started/welcome",
            "msapplication-config": "/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/browserconfig.xml",
            "og:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DGet%2BStarted%26title%3DWelcome%2Bto%2BOlostep%26description%3DOlostep%253A%2BInfrastructure%2Bfor%2Bthe%2BWeb%2527s%2Bsecond%2Buser.%2BThe%2Bbest%2Bsearch%252C%2Bscraping%2Band%2Bcrawling%2BAPI%2Bfor%2BAI.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "og:site_name": "Olostep Docs",
            "twitter:title": "Welcome to Olostep - Olostep Docs",
            "twitter:image:width": "1200",
            "generator": "Mintlify",
            "ogImage": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DGet%2BStarted%26title%3DWelcome%2Bto%2BOlostep%26description%3DOlostep%253A%2BInfrastructure%2Bfor%2Bthe%2BWeb%2527s%2Bsecond%2Buser.%2BThe%2Bbest%2Bsearch%252C%2Bscraping%2Band%2Bcrawling%2BAPI%2Bfor%2BAI.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "ogSiteName": "Olostep Docs",
            "twitter:card": "summary_large_image",
            "twitter:image:height": "630",
            "ogDescription": "Olostep: Infrastructure for the Web's second user. The best search, scraping and crawling API for AI.",
            "og:image:width": "1200",
            "favicon": "https://docs.olostep.com/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/android-chrome-192x192.png",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-a89408851a67",
            "sourceURL": "https://docs.olostep.com/get-started/welcome",
            "url": "https://docs.olostep.com/get-started/welcome",
            "statusCode": 200,
            "contentType": "text/html; charset=utf-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "861deaf0-07ac-4c57-9204-271ec3a38d39",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats",
          "title": "Difference between Fragementation in percent vs Page count in Index ...",
          "description": "Page count is the number of pages that the data in the table takes up, each page is 8kb. The only reliable way to decrease that is to delete ...",
          "position": 8,
          "markdown": "[Post reply](https://www.sqlservercentral.com/wp-login.php?redirect_to=https%3A%2F%2Fwww.sqlservercentral.com%2Fforums%2Ftopic%2Fdifference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats%23new-post)\n\n* * *\n\n# Difference between Fragementation in percent vs Page count in Index physical stats.\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 12:40 pm\n\n\n\n[#256845](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-256845)\n\n\n\n\n\n\n\nI would like to know about Fragementation in percent vs Page count. Whenever I tried to reorganize or rebuild the query, I see the following changes\n\n\n\nIndexTypeAvgPageFragmentationPageCounts\n\n\n\nCLUSTERED INDEX 66.6666666666667 3\n\n\n\nNONCLUSTERED INDEX 66.6666666666667 3\n\n\n\nHEAP63.7976346911958 12107\n\n\n\nCan someone explain me more detail way. If I further do any rebuild on clustered index, it remains same. Some times, the page count won't get reduced for example:\n\n\n\nIndexTypeAvgPageFragmentationPageCounts\n\n\n\nCLUSTERED INDEX 0.41958041958042 715\n\n\n\nWhat exactly this page count does and how can we reduce it or is it required to pay attention onto this page counts.\n\n- [SGT\\_squeequal](https://www.sqlservercentral.com/forums/user/sgtsqueequal)\n\n\n\n\n\n\n\n\n\nSSCertifiable\n\n\n\nPoints: 7167\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 1:29 pm\n\n\n\n[#1491343](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491343)\n\n\n\n\n\n\n\nHave a look at the stairways and indexing, there is plenty of information for you on there.\n\n\n\nas for pages, data is stored on pageses therefore if you have a page count of 10 then that is how many pages the data is held across\n\n\n\n\n\n\\*\\*\\*The first step is always the hardest \\*\\*\\*\\*\\*\\*\\*\n\n- [Gail Shaw](https://www.sqlservercentral.com/forums/user/gilamonster)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 1004485\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 1:56 pm\n\n\n\n[#1491361](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491361)\n\n\n\n\n\n\n\nPage count is the number of pages that the data in the table takes up, each page is 8kb. The only reliable way to decrease that is to delete data.\n\n\n\nDon't fuss over indexes with 3 pages, they won't defrag. The general recommendation is to worry about fragmentation once a table is over 1000 pages or so.\n\n\n\n\n\nGail Shaw\n\nMicrosoft Certified Master: SQL Server, MVP, M.Sc (Comp Sci)\n\n[**SQL In The Wild**](http://sqlinthewild.co.za/): Discussions on DB performance with occasional diversions into recoverability\n\n\n\n_We walk in the dark places no others will enter_\n\n_We stand on the bridge and no one may pass_\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:14 pm\n\n\n\n[#1491375](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491375)\n\n\n\n\n\n\n\nSo, if the table has more than 1000 pages, so how can we reduce it.\n\n- [Gail Shaw](https://www.sqlservercentral.com/forums/user/gilamonster)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 1004485\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:17 pm\n\n\n\n[#1491379](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491379)\n\n\n\n\n\n\n\nReduce what? Fragmentation or page count?\n\n\n\n\n\nGail Shaw\n\nMicrosoft Certified Master: SQL Server, MVP, M.Sc (Comp Sci)\n\n[**SQL In The Wild**](http://sqlinthewild.co.za/): Discussions on DB performance with occasional diversions into recoverability\n\n\n\n_We walk in the dark places no others will enter_\n\n_We stand on the bridge and no one may pass_\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:24 pm\n\n\n\n[#1491388](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491388)\n\n\n\n\n\n\n\nReduce page count\n\n- [Gail Shaw](https://www.sqlservercentral.com/forums/user/gilamonster)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 1004485\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:28 pm\n\n\n\n[#1491394](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491394)\n\n\n\n\n\n\n\nDelete data from the table. That's the only reliable way to reduce the size of the data in the table, which is what page count is.\n\n\n\nWhy are you fixated on page count? If a table has 8MB of data in it, it will have at least 1000 pages because 8MB of data, 8kb per page, 1000 pages.\n\n\n\n\n\nGail Shaw\n\nMicrosoft Certified Master: SQL Server, MVP, M.Sc (Comp Sci)\n\n[**SQL In The Wild**](http://sqlinthewild.co.za/): Discussions on DB performance with occasional diversions into recoverability\n\n\n\n_We walk in the dark places no others will enter_\n\n_We stand on the bridge and no one may pass_\n\n- [Lynn Pettis](https://www.sqlservercentral.com/forums/user/lynn-pettis)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 442462\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 2:28 pm\n\n\n\n[#1491395](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491395)\n\n\n\n\n\n\n\n\n\n\n\n\n\n> **DBA\\_SQL (5/22/2012)**\n>\n> * * *\n>\n> Reduce page count\n\n\n\n\n\n\n\n\n\nDelete data.\n\n\n\nTrying to understand why you are focusing on page count. As data is added to the database the page count is going to go up.\n\n- [DBA\\_Learner](https://www.sqlservercentral.com/forums/user/DBA_Learner)\n\n\n\n\n\n\n\n\n\nSSCarpal Tunnel\n\n\n\nPoints: 4228\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 3:04 pm\n\n\n\n[#1491433](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491433)\n\n\n\n\n\n\n\nI think I am confusing. So, if we have more data, we get more pages...and vice versa right...So, at this point i think it is good to focus on only fragmentation part rather than pages, because it all depends on data.\n\n- [Lynn Pettis](https://www.sqlservercentral.com/forums/user/lynn-pettis)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 442462\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 3:06 pm\n\n\n\n[#1491434](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491434)\n\n\n\n\n\n\n\nYes. As Gail said earlier, don't even worry about fragmentation until you hit about 1000 pages in the table.\n\n- [Jared](https://www.sqlservercentral.com/forums/user/sqlknowitall)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 61793\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nMay 22, 2012 at 3:06 pm\n\n\n\n[#1491436](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-1491436)\n\n\n\n\n\n\n\n\n\n\n\n\n\n> **DBA\\_SQL (5/22/2012)**\n>\n> * * *\n>\n> I think I am confusing. So, if we have more data, we get more pages...and vice versa right...So, at this point i think it is good to focus on only fragmentation part rather than pages, because it all depends on data.\n\n\n\n\n\n\n\n\n\nOk, but if the index is less than 1000 pages (or whatever you see fit), then don't worry about fragmentation.\n\n\n\n\n\nJared\n\nCE - Microsoft\n\n- [nmcquillen](https://www.sqlservercentral.com/forums/user/nmcquillen)\n\n\n\n\n\n\n\n\n\nGrasshopper\n\n\n\nPoints: 11\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nAugust 13, 2018 at 1:55 pm\n\n\n\n[#2001552](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-2001552)\n\n\n\n\n\n\n\nI realize this is a very old post, but a few things.\u00a0 If your index is getting scanned then ignore that 1000 pages business.\u00a0 I've seen scans on indexes choke with around a 300 page\\_count, so check your index scan stats.\u00a0 If you don't think you should be having scans, make sure whatever operation/query is sargable, no implicit\\_conversion, etc.\u00a0 Also if you have a higher page\\_count than you think should be required (considering offset/header) you might be dealing splitting while updating data bigger than the original slot so your page density isn't full when that occurs.\u00a0 Rebuild operations should handle this or you could go full tilt and take the db offline, defrag the mdf (to reduce physical fragmentation by hopefully getting contiguous mapping), rebuild indexes (sort in tempdb=on maxdop=1 for serial builds done on a separate disk than destination), and possibly mess around with fillfactor to mitigate splits.\n\n- [ScottPletcher](https://www.sqlservercentral.com/forums/user/scottpletcher)\n\n\n\n\n\n\n\n\n\nSSC Guru\n\n\n\nPoints: 101249\n\n\n\n\n\n[More actions](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#)\n\n\n\n\n\n\n\n\n\n\n\n\n\nAugust 13, 2018 at 2:02 pm\n\n\n\n[#2001553](https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats#post-2001553)\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n> **nmcquillen - Monday, August 13, 2018 1:55 PM**\n>\n> I realize this is a very old post, but a few things.\u00a0 If your index is getting scanned then ignore that 1000 pages business.\u00a0 I've seen scans on indexes choke with around a 300 page\\_count, so check your index scan stats.\u00a0 If you don't think you should be having scans, make sure whatever operation/query is sargable, no implicit\\_conversion, etc.\u00a0 Also if you have a higher page\\_count than you think should be required (considering offset/header) you might be dealing splitting while updating data bigger than the original slot so your page density isn't full when that occurs.\u00a0 Rebuild operations should handle this or you could go full tilt and take the db offline, defrag the mdf (to reduce physical fragmentation by hopefully getting contiguous mapping), rebuild indexes (sort in tempdb=on maxdop=1 for serial builds done on a separate disk than destination), and possibly mess around with fillfactor to mitigate splits.\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\nAbsolutely.\u00a0 Even the person who originally came up with the \"1000 pages\" admitted it was just a round number, with no real analysis to it.\u00a0 Determine the best clustering index and apply it, no matter how many rows the table (currently) has.\u00a0 You can also force a table(s) with less than 8 total pages to be put into a single extent, which for a busy, small table can reduce I/O.\n\n\n\nAs to ways to reduce data size, you have other options besides just deleting data (which usually isn't a viable business option at all);\n\n1) if you have an edition of SQL that supports it, compress the data.\n\n2) if you don't, encode data, esp. character data.\u00a0 That is, use a numeric id/code in place of a longer string value.\n\nThere are some others, but those should give you the biggest payback for the least overall effort.\n\n\n\n\n\n\n\nSQL DBA,SQL Server MVP(07, 08, 09) \"It's a dog-eat-dog world, and I'm wearing Milk-Bone underwear.\" \"Norm\", on \"Cheers\". Also from \"Cheers\", from \"Carla\": \"You need to know 3 things about Tortelli men: Tortelli men draw women like flies; Tortelli men treat women like flies; Tortelli men's brains are in their flies\".\n\n\nViewing 13 posts - 1 through 13 (of 13 total)\n\nYou must be logged in to reply to this topic. [Login to reply](https://www.sqlservercentral.com/wp-login.php?redirect_to=https%3A%2F%2Fwww.sqlservercentral.com%2Fforums%2Ftopic%2Fdifference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats%23new-post)",
          "metadata": {
            "og:type": "article",
            "robots": "index, follow, max-image-preview:large, max-snippet:-1, max-video-preview:-1",
            "viewport": "width=device-width, initial-scale=1",
            "twitter:site": "@sqlservercentrl",
            "generator": "WordPress 6.8.1",
            "ogLocale": "en_GB",
            "ogSiteName": "SQLServerCentral",
            "ahrefs-site-verification": "b21134b1225ebf149e1266990fcc7b2bb3d9812fb4c9cdb1db0e5b103b2be6d2",
            "og:locale": "en_GB",
            "ogUrl": "https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats",
            "twitter:card": "summary_large_image",
            "ogDescription": "Difference between Fragementation in percent vs Page count in Index physical stats. Forum \u2013 Learn more on SQLServerCentral",
            "wp:id": "256845",
            "title": "Difference between Fragementation in percent vs Page count in Index physical stats. \u2013 SQLServerCentral Forums",
            "ogTitle": "Difference between Fragementation in percent vs Page count in Index physical stats. \u2013 SQLServerCentral Forums",
            "projectnami:version": "3.8.1",
            "og:site_name": "SQLServerCentral",
            "description": "Difference between Fragementation in percent vs Page count in Index physical stats. Forum \u2013 Learn more on SQLServerCentral",
            "og:description": "Difference between Fragementation in percent vs Page count in Index physical stats. Forum \u2013 Learn more on SQLServerCentral",
            "language": "en-GB",
            "og:url": "https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats",
            "og:title": "Difference between Fragementation in percent vs Page count in Index physical stats. \u2013 SQLServerCentral Forums",
            "favicon": "https://www.sqlservercentral.com/wp-content/uploads/2019/04/favicon.ico",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-ae2dd22edbed",
            "sourceURL": "https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats",
            "url": "https://www.sqlservercentral.com/forums/topic/difference-between-fragementation-in-percent-vs-page-count-in-index-physical-stats",
            "statusCode": 200,
            "contentType": "text/html; charset=UTF-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "6f81ae12-42af-48a9-9122-ff3e5275065f",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17",
          "title": "Estimate the Size of a Clustered Index - SQL Server | Microsoft Learn",
          "description": "The first index level above the leaf level stores 1,000 index rows, which is one index row per leaf page, and 25 index rows can fit per page.",
          "position": 9,
          "markdown": "Table of contents Exit editor mode\n\nAsk LearnAsk Learn\n\nReading modeTable of contents[Read in English](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17)Add to CollectionsAdd to Plans[Edit](https://github.com/MicrosoftDocs/sql-docs/blob/live/docs/relational-databases/databases/estimate-the-size-of-a-clustered-index.md)\n\n* * *\n\nCopy MarkdownPrint\n\n* * *\n\nNote\n\nAccess to this page requires authorization. You can try [signing in](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#) or changing directories.\n\n\nAccess to this page requires authorization. You can try changing directories.\n\n\n# Estimate the size of a clustered index\n\nFeedback\n\nSummarize this article for me\n\n\n**Applies to:**![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [SQL Server](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [Azure SQL Database](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [Azure SQL Managed Instance](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)![](https://learn.microsoft.com/en-us/sql/includes/media/yes-icon.svg?view=sql-server-ver17) [SQL database in Microsoft Fabric](https://learn.microsoft.com/en-us/sql/sql-server/sql-docs-navigation-guide?view=sql-server-ver17#applies-to)\n\nYou can use the following steps to estimate the amount of space that is required to store data in a clustered index:\n\n1. Calculate the space used to store data in the leaf level of the clustered index.\n2. Calculate the space used to store index information for the clustered index.\n3. Total the calculated values.\n\n[Section titled: Step 1. Calculate the space used to store data in the leaf level](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#step-1-calculate-the-space-used-to-store-data-in-the-leaf-level)\n\n## Step 1. Calculate the space used to store data in the leaf level\n\n01. Specify the number of rows that are present in the table:\n\n    - _**Num\\_Rows**_ = number of rows in the table\n02. Specify the number of fixed-length and variable-length columns and calculate the space that is required for their storage:\n\n    Calculate the space that each of these groups of columns occupies within the data row. The size of a column depends on the data type and length specification.\n\n    - _**Num\\_Cols**_ = total number of columns (fixed-length and variable-length)\n    - _**Fixed\\_Data\\_Size**_ = total byte size of all fixed-length columns\n    - _**Num\\_Variable\\_Cols**_ = number of variable-length columns\n    - _**Max\\_Var\\_Size**_ = maximum byte size of all variable-length columns\n03. If the clustered index is nonunique, account for the _uniqueifier_ column:\n\n    The uniqueifier is a nullable, variable-length column. It's non-null and 4 bytes in size in rows that have nonunique key values. This value is part of the index key and is required to make sure that every row has a unique key value.\n\n\n    - _**Num\\_Cols**_ = _**Num\\_Cols**_ \\+ 1\n    - _**Num\\_Variable\\_Cols**_ = _**Num\\_Variable\\_Cols**_ \\+ 1\n    - _**Max\\_Var\\_Size**_ = _**Max\\_Var\\_Size**_ \\+ 4\n\nThese modifications assume that all values are nonunique.\n\n04. Part of the row, known as the null bitmap, is reserved to manage column nullability. Calculate its size:\n\n\n    - _**Null\\_Bitmap**_ = 2 + (( _**Num\\_Cols**_ \\+ 7) / 8)\n\nOnly the integer part of the previous expression should be used; discard any remainder.\n\n05. Calculate the variable-length data size:\n\n    If there are variable-length columns in the table, determine how much space is used to store the columns within the row:\n\n\n    - _**Variable\\_Data\\_Size**_ = 2 + ( _**Num\\_Variable\\_Cols**_ x 2) + _**Max\\_Var\\_Size**_\n\nThe bytes added to _**Max\\_Var\\_Size**_ are for tracking each variable column. This formula assumes that all variable-length columns are 100 percent full. If you anticipate that a smaller percentage of the variable-length column storage space will be used, you can adjust the _**Max\\_Var\\_Size**_ value by that percentage to yield a more accurate estimate of the overall table size.\n\nYou can combine **varchar**, **nvarchar**, **varbinary**, or **sql\\_variant** columns that cause the total defined table width to exceed 8,060 bytes. The length of each one of these columns must still fall within the limit of 8,000 bytes for a **varchar**, **varbinary**, or **sql\\_variant** column, and 4,000 bytes for **nvarchar** columns. However, their combined widths might exceed the 8,060-byte limit in a table.\n\nIf there are no variable-length columns, set _**Variable\\_Data\\_Size**_ to 0.\n\n06. Calculate the total row size:\n\n\n    - _**Row\\_Size**_ = _**Fixed\\_Data\\_Size**_ \\+ _**Variable\\_Data\\_Size**_ \\+ _**Null\\_Bitmap**_ \\+ 4\n\nThe value 4 is the row header overhead of a data row.\n\n07. Calculate the number of rows per page (8,096 free bytes per page):\n\n\n    - _**Rows\\_Per\\_Page**_ = 8096 / ( _**Row\\_Size**_ \\+ 2)\n\nBecause rows don't span pages, the number of rows per page should be rounded down to the nearest whole row. The value 2 in the formula is for the row's entry in the slot array of the page.\n\n08. Calculate the number of reserved free rows per page, based on the [fill factor](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/specify-fill-factor-for-an-index?view=sql-server-ver17) specified:\n\n\n    - _**Free\\_Rows\\_Per\\_Page**_ = 8096 x ((100 - _**Fill\\_Factor**_) / 100) / ( _**Row\\_Size**_ \\+ 2)\n\nThe fill factor used in the calculation is an integer value instead of a percentage. Because rows don't span pages, the number of rows per page should be rounded down to the nearest whole row. As the fill factor grows, more data is stored on each page and there are fewer pages. The value 2 in the formula is for the row's entry in the slot array of the page.\n\n09. Calculate the number of pages required to store all the rows:\n\n\n    - _**Num\\_Leaf\\_Pages**_ = _**Num\\_Rows**_ / ( _**Rows\\_Per\\_Page**_ \\- _**Free\\_Rows\\_Per\\_Page**_)\n\nThe number of pages estimated should be rounded up to the nearest whole page.\n\n10. Calculate the amount of space that is required to store the data in the leaf level (8,192 total bytes per page):\n\n    - _**Leaf\\_space\\_used**_ = 8192 x _**Num\\_Leaf\\_Pages**_\n\n[Section titled: Step 2. Calculate the space used to store index information](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#step-2-calculate-the-space-used-to-store-index-information)\n\n## Step 2. Calculate the space used to store index information\n\nYou can use the following steps to estimate the amount of space that is required to store the upper levels of the index:\n\n1. Specify the number of fixed-length and variable-length columns in the index key and calculate the space that is required for their storage:\n\nThe key columns of an index can include fixed-length and variable-length columns. To estimate the interior level index row size, calculate the space that each of these groups of columns occupies within the index row. The size of a column depends on the data type and length specification.\n\n   - _**Num\\_Key\\_Cols**_ = total number of key columns (fixed-length and variable-length)\n   - _**Fixed\\_Key\\_Size**_ = total byte size of all fixed-length key columns\n   - _**Num\\_Variable\\_Key\\_Cols**_ = number of variable-length key columns\n   - _**Max\\_Var\\_Key\\_Size**_ = maximum byte size of all variable-length key columns\n2. Account for any uniqueifier needed if the index is nonunique:\n\nThe uniqueifier is a nullable, variable-length column. It's non-null and 4 bytes in size in rows that have nonunique index key values. This value is part of the index key and is required to make sure that every row has a unique key value.\n\n\n   - _**Num\\_Key\\_Cols**_ = _**Num\\_Key\\_Cols**_ \\+ 1\n   - _**Num\\_Variable\\_Key\\_Cols**_ = _**Num\\_Variable\\_Key\\_Cols**_ \\+ 1\n   - _**Max\\_Var\\_Key\\_Size**_ = _**Max\\_Var\\_Key\\_Size**_ \\+ 4\n\nThese modifications assume that all values are nonunique.\n\n3. Calculate the null bitmap size:\n\nIf there are nullable columns in the index key, part of the index row is reserved for the null bitmap. Calculate its size:\n\n\n   - _**Index\\_Null\\_Bitmap**_ = 2 + ((number of columns in the index row + 7) / 8)\n\nOnly the integer part of the previous expression should be used. Discard any remainder.\n\nIf there are no nullable key columns, set _**Index\\_Null\\_Bitmap**_ to 0.\n\n4. Calculate the variable-length data size:\n\nIf there are variable-length columns in the index, determine how much space is used to store the columns within the index row:\n\n\n   - _**Variable\\_Key\\_Size**_ = 2 + ( _**Num\\_Variable\\_Key\\_Cols**_ x 2) + _**Max\\_Var\\_Key\\_Size**_\n\nThe bytes added to _**Max\\_Var\\_Key\\_Size**_ are for tracking each variable-length column. This formula assumes that all variable-length columns are 100 percent full. If you anticipate that a smaller percentage of the variable-length column storage space will be used, you can adjust the _**Max\\_Var\\_Key\\_Size**_ value by that percentage to yield a more accurate estimate of the overall table size.\n\nIf there are no variable-length columns, set _**Variable\\_Key\\_Size**_ to 0.\n\n5. Calculate the index row size:\n\n   - _**Index\\_Row\\_Size**_ = _**Fixed\\_Key\\_Size**_ \\+ _**Variable\\_Key\\_Size**_ \\+ _**Index\\_Null\\_Bitmap**_ \\+ 1 (for row header overhead of an index row) + 6 (for the child page ID pointer)\n6. Calculate the number of index rows per page (8,096 free bytes per page):\n\n\n   - _**Index\\_Rows\\_Per\\_Page**_ = 8096 / ( _**Index\\_Row\\_Size**_ \\+ 2)\n\nBecause index rows don't span pages, the number of index rows per page should be rounded down to the nearest whole row. The `2` in the formula is for the row's entry in the page's slot array.\n\n7. Calculate the number of levels in the index:\n\n\n   - _**Non-leaf\\_Levels**_ = 1 + log (Index\\_Rows\\_Per\\_Page) ( _**Num\\_Leaf\\_Pages**_ / _**Index\\_Rows\\_Per\\_Page**_)\n\nRound this value up to the nearest whole number. This value doesn't include the leaf level of the clustered index.\n\n8. Calculate the number of nonleaf pages in the index:\n\n\n   - _**Num\\_Index\\_Pages =**_ \u2211Level ( _**Num\\_Leaf\\_Pages**_ / ( _**Index\\_Rows\\_Per\\_Page**_ ^ _**Level**_))\n\n     where 1 <= Level <= _**Non-leaf\\_Levels**_\n\n\nRound each summand up to the nearest whole number. As a simple example, consider an index where _**Num\\_Leaf\\_Pages**_ = 1000 and _**Index\\_Rows\\_Per\\_Page**_ = 25\\. The first index level above the leaf level stores 1,000 index rows, which is one index row per leaf page, and 25 index rows can fit per page. This means that 40 pages are required to store those 1,000 index rows. The next level of the index has to store 40 rows. This means it requires two pages. The final level of the index has to store two rows. This means it requires one page. This gives 43 nonleaf index pages. When these numbers are used in the previous formulas, the outcome is as follows:\n\n   - _**Non-leaf\\_Levels**_ = 1 + log(25) (1000 / 25) = 3\n\n   - _**Num\\_Index\\_Pages**_ = 1000/(25^3)+ 1000/(25^2) + 1000/(25^1) = 1 + 2 + 40 = 43, which is the number of pages described in the example.\n9. Calculate the size of the index (8,192 total bytes per page):\n\n   - _**Index\\_Space\\_Used**_ = 8192 x _**Num\\_Index\\_Pages**_\n\n[Section titled: Step 3. Total the calculated values](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#step-3-total-the-calculated-values)\n\n## Step 3. Total the calculated values\n\nTotal the values obtained from the previous two steps:\n\n- Clustered index size (bytes) = _**Leaf\\_Space\\_Used**_ \\+ _**Index\\_Space\\_used**_\n\nThis calculation doesn't consider the following conditions:\n\n- **Partitioning**: The space overhead from partitioning is minimal, but complex to calculate. It isn't important to include.\n\n- **Allocation pages**: There's at least one IAM page used to track the pages allocated to a heap. The space overhead is minimal, and there's no algorithm to deterministically calculate exactly how many IAM pages will be used.\n\n- **Large object (LOB) values**: The algorithm to determine exactly how much space will be used to store the LOB data types **varchar(max)**, **varbinary(max)**, **nvarchar(max)**, **text**, **ntext**, **xml**, and **image** values is complex. It's sufficient to just add the average size of the LOB values that are expected, multiply by _**Num\\_Rows**_, and add that to the total clustered index size.\n\n- **Compression**: You can't precalculate the size of a compressed index.\n\n- **Sparse columns**: For information about the space requirements of sparse columns, see [Use sparse columns](https://learn.microsoft.com/en-us/sql/relational-databases/tables/use-sparse-columns?view=sql-server-ver17).\n\n\n[Section titled: Related content](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#related-content)\n\n## Related content\n\n- [Clustered and nonclustered indexes](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/clustered-and-nonclustered-indexes-described?view=sql-server-ver17)\n- [Estimate the size of a table](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-table?view=sql-server-ver17)\n- [Create a clustered index](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/create-clustered-indexes?view=sql-server-ver17)\n- [Create nonclustered indexes](https://learn.microsoft.com/en-us/sql/relational-databases/indexes/create-nonclustered-indexes?view=sql-server-ver17)\n- [Estimate the size of a nonclustered index](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-nonclustered-index?view=sql-server-ver17)\n- [Estimate the size of a heap](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-heap?view=sql-server-ver17)\n- [Estimate the size of a database](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-database?view=sql-server-ver17)\n\nReading mode disabled\n\n* * *\n\n## Feedback\n\nWas this page helpful?\n\n\nYesNoNo\n\nNeed help with this topic?\n\n\nWant to try using Ask Learn to clarify or guide you through this topic?\n\n\nAsk LearnAsk Learn\n\nSuggest a fix?\n\n* * *\n\n## Additional resources\n\n* * *\n\n- Last updated on 07/20/2026\n\nAsk Learn is an AI assistant that can answer questions, clarify concepts, and define terms using trusted Microsoft documentation.\n\nPlease sign in to use Ask Learn.\n\n[Sign in](https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17#)",
          "metadata": {
            "language": "en-us",
            "asset_id": "relational-databases/databases/estimate-the-size-of-a-clustered-index",
            "moniker_range_name": "2bbf5b3331aa22b1ac18ea8893c72c8a",
            "monikerRange": "=azuresqldb-current || >=sql-server-2017 || =fabric-sqldb",
            "scope": "sql",
            "feedback_help_link_type": "get-help-at-qna",
            "site_name": "Docs",
            "og:image": "https://learn.microsoft.com/en-us/media/open-graph-image.png",
            "og:url": "https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17",
            "locale": "en-us",
            "og:description": "Use this procedure to estimate the amount of space that is required to store data in a clustered index in SQL Server.",
            "ogImage": "https://learn.microsoft.com/en-us/media/open-graph-image.png",
            "ms.reviewer": "randolphwest",
            "platform_id": "a873750e-a101-e19b-866c-65e1bbfd4cc5",
            "ms.update-cycle": "1825-days",
            "recommendations": "true",
            "original_content_git_url": "https://github.com/MicrosoftDocs/sql-docs-pr/blob/live/docs/relational-databases/databases/estimate-the-size-of-a-clustered-index.md",
            "ogDescription": "Use this procedure to estimate the amount of space that is required to store data in a clustered index in SQL Server.",
            "cmProducts": [
              "https://authoring-docs-microsoft.poolparty.biz/devrel/cbe4ca68-43ac-4375-aba5-5945a6394c20",
              "https://authoring-docs-microsoft.poolparty.biz/devrel/6ab7faaf-d791-4a26-96a2-3b11738538e7",
              "https://authoring-docs-microsoft.poolparty.biz/devrel/8b896464-3b7d-4e1f-84b0-9bb45aeb5f64"
            ],
            "spProducts": [
              "https://authoring-docs-microsoft.poolparty.biz/devrel/ced846cc-6a3c-4c8f-9dfb-3de0e90e2742",
              "https://authoring-docs-microsoft.poolparty.biz/devrel/302e28b0-1f09-4811-9a9b-2a72e0770581",
              "https://authoring-docs-microsoft.poolparty.biz/devrel/b1d2d671-9549-46e8-918c-24349120dbf5"
            ],
            "toc_rel": "../../toc.json",
            "author": "WilliamDAssafMSFT",
            "feedback_system": "Standard",
            "ms.custom": "ignite-2025",
            "viewport": "width=device-width, initial-scale=1.0",
            "page_type": "conceptual",
            "git_commit_id": "6195ae0be17d3691eed9bd682098620c5102fd4e",
            "config_moniker_range": "=azuresqldb-current || =azuresqldb-mi-current || =azure-sqldw-latest || >=aps-pdw-2016 || >=sql-server-2017 || >=sql-server-linux-2017 || =fabric || =fabric-sqldb",
            "depot_name": "SQL.sql-content",
            "color-scheme": "light dark",
            "twitter:card": "summary_large_image",
            "og:type": "website",
            "og:title": "Estimate the Size of a Clustered Index - SQL Server",
            "item_type": "Content",
            "title": "Estimate the Size of a Clustered Index - SQL Server | Microsoft Learn",
            "ms.author": "wiassaf",
            "ms.subservice": "supportability",
            "default_moniker": "sql-server-ver17",
            "description": "Use this procedure to estimate the amount of space that is required to store data in a clustered index in SQL Server.",
            "ms.date": "2025-09-22T00:00:00Z",
            "monikers": [
              "azuresqldb-current",
              "fabric-sqldb",
              "sql-server-2017",
              "sql-server-ver15",
              "sql-server-ver16",
              "sql-server-ver17"
            ],
            "feedback_help_link_url": "https://learn.microsoft.com/answers/tags/191/sql-server",
            "twitter:site": "@MicrosoftLearn",
            "ms.service": "sql",
            "ogTitle": "Estimate the Size of a Clustered Index - SQL Server",
            "word_count": "1613",
            "markdown_url": "https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17&accept=text/markdown",
            "previous_tlsh_hash": "5B7C0B12742CD702DE825E1B215FD89434F09E0AB6A26EA8992F3932625A2C730FAEE467C6336FD027353D637345B1EDA7D55F6A409803B422E2282D451D8285D75C4BFFF4",
            "toc_preview": "true",
            "document_version_independent_id": "64aa8ace-0107-570a-1b6e-2c17505a2081",
            "document_id": "37903150-ad04-bdde-e756-adb6ba228cec",
            "uhfHeaderId": "MSDocsHeader-DocsSQL",
            "gitcommit": "https://github.com/MicrosoftDocs/sql-docs-pr/blob/6195ae0be17d3691eed9bd682098620c5102fd4e/docs/relational-databases/databases/estimate-the-size-of-a-clustered-index.md",
            "ms.topic": "how-to",
            "pdf_url_template": "https://learn.microsoft.com/pdfstore/en-us/SQL.sql-content/{branchName}{pdfName}",
            "schema": "Conceptual",
            "github_feedback_content_git_url": "https://github.com/MicrosoftDocs/sql-docs/blob/live/docs/relational-databases/databases/estimate-the-size-of-a-clustered-index.md",
            "og:image:alt": "Microsoft Learn",
            "updated_at": "2026-07-20T22:35:00Z",
            "breadcrumb_path": "../../breadcrumb/toc.json",
            "feedback_product_url": "https://feedback.azure.com/d365community/forum/04fe6ee0-3b25-ec11-b6e6-000d3a4f0da0",
            "ogUrl": "https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17",
            "source_path": "docs/relational-databases/databases/estimate-the-size-of-a-clustered-index.md",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-b1466347b5f3",
            "sourceURL": "https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17",
            "url": "https://learn.microsoft.com/en-us/sql/relational-databases/databases/estimate-the-size-of-a-clustered-index?view=sql-server-ver17",
            "statusCode": 200,
            "contentType": "text/html",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "cd4c94d3-9b4c-4879-aa9f-0bf91ff71e59",
            "creditsUsed": 1
          }
        },
        {
          "url": "https://docs.olostep.com/api-reference/scrapes/create",
          "title": "Create Scrape - Olostep Docs",
          "description": "## Documentation Index\n```\ncurl --request POST \\\n  --url https://api.olostep.com/v1/scrapes \\\n  --header 'Authorization: Bearer <token>' \\\n  --header 'Content-Type: application/json' \\\n  --data '\n{\n  \"url_to_scrape\": \"<string>\",\n  \"wait_before_scraping\": 123,\n  \"formats\": [],\n  \"actions\": [\\\n    {\\\n      \"type\": \"wait\",\\\n      \"milliseconds\": 1\\\n    }\\\n  ],\n  \"country\": \"<string>\",\n  \"remove_images\": false,\n  \"remove_class_names\": [\\\n    \"<string>\"\\\n  ],\n  \"llm_extract\": {\n    \"schema\": {}\n  },\n  \"links_on_page\": {\n    \"query_to_order_links_by\": \"<string>\",\n    \"include_links\": [\\\n      \"<string>\"\\\n    ],\n    \"exclude_links\": [\\\n      \"<string>\"\\\n    ]\n  },\n  \"screen_size\": {\n    \"screen_width\": 123,\n    \"screen_height\": 123\n  },\n  \"screenshot\": {\n    \"full_page\": true\n  },\n  \"metadata\": {},\n  \"max_age\": 0\n}\n'\n```\n\n```\nimport requests\n\nurl = \"https://api.olostep.com/v1/scrapes\"\n\npayload = {\n    \"url_to_scrape\": \"<string>\",\n    \"wait_before_scraping\": 123,\n    \"formats\": [],\n    \"actions\": [\\\n        {\\\n            \"type\": \"wait\",\\\n            \"milliseconds\": 1\\\n        }\\\n    ],\n    \"country\": \"<string>\",\n    \"remove_images\": False,\n    \"remove_class_names\": [\"<string>\"],\n    \"llm_extract\": { \"schema\": {} },\n    \"links_on_page\": {\n        \"query_to_order_links_by\": \"<string>\",\n        \"include_links\": [\"<string>\"],\n        \"exclude_links\": [\"<string>\"]\n    },\n    \"screen_size\": {\n        \"screen_width\": 123,\n        \"screen_height\": 123\n    },\n    \"screenshot\": { \"full_page\": True },\n    \"metadata\": {},\n    \"max_age\": 0\n}\nheaders = {\n    \"Authorization\": \"Bearer <token>\",\n    \"Content-Type\": \"application/json\"\n}\n\nresponse = requests.post(url, json=payload, headers=headers)\n\nprint(response.text)\n```\n\n```\nconst options = {\n  method: 'POST',\n  headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},\n  body: JSON.stringify({\n    url_to_scrape: '<string>',\n    wait_before_scraping: 123,\n    formats: [],\n    actions: [{type: 'wait', milliseconds: 1}],\n    country: '<string>',\n    remove_images: false,\n    remove_class_names: ['<string>'],\n    llm_extract: {schema: {}},\n    links_on_page: {\n      query_to_order_links_by: '<string>',\n      include_links: ['<string>'],\n      exclude_links: ['<string>']\n    },\n    screen_size: {screen_width: 123, screen_height: 123},\n    screenshot: {full_page: true},\n    metadata: {},\n    max_age: 0\n  })\n};\n\nfetch('https://api.olostep.com/v1/scrapes', options)\n  .then(res => res.json())\n  .then(res => console.log(res))\n  .catch(err => console.error(err));\n```\n\n```\n<?php\n\n$curl = curl_init();\n\ncurl_setopt_array($curl, [\\\n  CURLOPT_URL => \"https://api.olostep.com/v1/scrapes\",\\\n  CURLOPT_RETURNTRANSFER => true,\\\n  CURLOPT_ENCODING => \"\",\\\n  CURLOPT_MAXREDIRS => 10,\\\n  CURLOPT_TIMEOUT => 30,\\\n  CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,\\\n  CURLOPT_CUSTOMREQUEST => \"POST\",\\\n  CURLOPT_POSTFIELDS => json_encode([\\\n    'url_to_scrape' => '<string>',\\\n    'wait_before_scraping' => 123,\\\n    'formats' => [\\\n\\\n    ],\\\n    'actions' => [\\\n        [\\\n                'type' => 'wait',\\\n                'milliseconds' => 1\\\n        ]\\\n    ],\\\n    'country' => '<string>',\\\n    'remove_images' => false,\\\n    'remove_class_names' => [\\\n        '<string>'\\\n    ],\\\n    'llm_extract' => [\\\n        'schema' => [\\\n\\\n        ]\\\n    ],\\\n    'links_on_page' => [\\\n        'query_to_order_links_by' => '<string>',\\\n        'include_links' => [\\\n                '<string>'\\\n        ],\\\n        'exclude_links' => [\\\n                '<string>'\\\n        ]\\\n    ],\\\n    'screen_size' => [\\\n        'screen_width' => 123,\\\n        'screen_height' => 123\\\n    ],\\\n    'screenshot' => [\\\n        'full_page' => true\\\n    ],\\\n    'metadata' => [\\\n\\\n    ],\\\n    'max_age' => 0\\\n  ]),\\\n  CURLOPT_HTTPHEADER => [\\\n    \"Authorization: Bearer <token>\",\\\n    \"Content-Type: application/json\"\\\n  ],\\\n]);\n\n$response = curl_exec($curl);\n$err = curl_error($curl);\n\ncurl_close($curl);\n\nif ($err) {\n  echo \"cURL Error #:\" . $err;\n} else {\n  echo $response;\n}\n```\n\n```\ncurl --request POST \\\n  --url https://api.olostep.com/v1/scrapes \\\n  --header 'Authorization: Bearer <token>' \\\n  --header 'Content-Type: application/json' \\\n  --data '\n{\n  \"url_to_scrape\": \"<string>\",\n  \"wait_before_scraping\": 123,\n  \"formats\": [],\n  \"actions\": [\\\n    {\\\n      \"type\": \"wait\",\\\n      \"milliseconds\": 1\\\n    }\\\n  ],\n  \"country\": \"<string>\",\n  \"remove_images\": false,\n  \"remove_class_names\": [\\\n    \"<string>\"\\\n  ],\n  \"llm_extract\": {\n    \"schema\": {}\n  },\n  \"links_on_page\": {\n    \"query_to_order_links_by\": \"<string>\",\n    \"include_links\": [\\\n      \"<string>\"\\\n    ],\n    \"exclude_links\": [\\\n      \"<string>\"\\\n    ]\n  },\n  \"screen_size\": {\n    \"screen_width\": 123,\n    \"screen_height\": 123\n  },\n  \"screenshot\": {\n    \"full_page\": true\n  },\n  \"metadata\": {},\n  \"max_age\": 0\n}\n'\n```\n\n```\nimport requests\n\nurl = \"https://api.olostep.com/v1/scrapes\"\n\npayload = {\n    \"url_to_scrape\": \"<string>\",\n    \"wait_before_scraping\": 123,\n    \"formats\": [],\n    \"actions\": [\\\n        {\\\n            \"type\": \"wait\",\\\n            \"milliseconds\": 1\\\n        }\\\n    ],\n    \"country\": \"<string>\",\n    \"remove_images\": False,\n    \"remove_class_names\": [\"<string>\"],\n    \"llm_extract\": { \"schema\": {} },\n    \"links_on_page\": {\n        \"query_to_order_links_by\": \"<string>\",\n        \"include_links\": [\"<string>\"],\n        \"exclude_links\": [\"<string>\"]\n    },\n    \"screen_size\": {\n        \"screen_width\": 123,\n        \"screen_height\": 123\n    },\n    \"screenshot\": { \"full_page\": True },\n    \"metadata\": {},\n    \"max_age\": 0\n}\nheaders = {\n    \"Authorization\": \"Bearer <token>\",\n    \"Content-Type\": \"application/json\"\n}\n\nresponse = requests.post(url, json=payload, headers=headers)\n\nprint(response.text)\n```\n\n```\nconst options = {\n  method: 'POST',\n  headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},\n  body: JSON.stringify({\n    url_to_scrape: '<string>',\n    wait_before_scraping: 123,\n    formats: [],\n    actions: [{type: 'wait', milliseconds: 1}],\n    country: '<string>',\n    remove_images: false,\n    remove_class_names: ['<string>'],\n    llm_extract: {schema: {}},\n    links_on_page: {\n      query_to_order_links_by: '<string>',\n      include_links: ['<string>'],\n      exclude_links: ['<string>']\n    },\n    screen_size: {screen_width: 123, screen_height: 123},\n    screenshot: {full_page: true},\n    metadata: {},\n    max_age: 0\n  })\n};\n\nfetch('https://api.olostep.com/v1/scrapes', options)\n  .then(res => res.json())\n  .then(res => console.log(res))\n  .catch(err => console.error(err));\n```\n\n```\n<?php\n\n$curl = curl_init();\n\ncurl_setopt_array($curl, [\\\n  CURLOPT_URL => \"https://api.olostep.com/v1/scrapes\",\\\n  CURLOPT_RETURNTRANSFER => true,\\\n  CURLOPT_ENCODING => \"\",\\\n  CURLOPT_MAXREDIRS => 10,\\\n  CURLOPT_TIMEOUT => 30,\\\n  CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,\\\n  CURLOPT_CUSTOMREQUEST => \"POST\",\\\n  CURLOPT_POSTFIELDS => json_encode([\\\n    'url_to_scrape' => '<string>',\\\n    'wait_before_scraping' => 123,\\\n    'formats' => [\\\n\\\n    ],\\\n    'actions' => [\\\n        [\\\n                'type' => 'wait',\\\n                'milliseconds' => 1\\\n        ]\\\n    ],\\\n    'country' => '<string>',\\\n    'remove_images' => false,\\\n    'remove_class_names' => [\\\n        '<string>'\\\n    ],\\\n    'llm_extract' => [\\\n        'schema' => [\\\n\\\n        ]\\\n    ],\\\n    'links_on_page' => [\\\n        'query_to_order_links_by' => '<string>',\\\n        'include_links' => [\\\n                '<string>'\\\n        ],\\\n        'exclude_links' => [\\\n                '<string>'\\\n        ]\\\n    ],\\\n    'screen_size' => [\\\n        'screen_width' => 123,\\\n        'screen_height' => 123\\\n    ],\\\n    'screenshot' => [\\\n        'full_page' => true\\\n    ],\\\n    'metadata' => [\\\n\\\n    ],\\\n    'max_age' => 0\\\n  ]),\\\n  CURLOPT_HTTPHEADER => [\\\n    \"Authorization: Bearer <token>\",\\\n    \"Content-Type: application/json\"\\\n  ],\\\n]);\n\n$response = curl_exec($curl);\n$err = curl_error($curl);\n\ncurl_close($curl);\n\nif ($err) {\n  echo \"cURL Error #:\" . $err;\n} else {\n  echo $response;\n}\n```",
          "position": 10,
          "markdown": "> ## Documentation Index\n>\n> Fetch the complete documentation index at: [/llms.txt](https://docs.olostep.com/llms.txt)\n>\n> Use this file to discover all available pages before exploring further.\n\n[Skip to main content](https://docs.olostep.com/api-reference/scrapes/create#content-area)\n\nInitiate a web page scrape\n\ncURL\n\n```\ncurl --request POST \\\n  --url https://api.olostep.com/v1/scrapes \\\n  --header 'Authorization: Bearer <token>' \\\n  --header 'Content-Type: application/json' \\\n  --data '\n{\n  \"url_to_scrape\": \"<string>\",\n  \"wait_before_scraping\": 123,\n  \"formats\": [],\n  \"actions\": [\\\n    {\\\n      \"type\": \"wait\",\\\n      \"milliseconds\": 1\\\n    }\\\n  ],\n  \"country\": \"<string>\",\n  \"remove_images\": false,\n  \"remove_class_names\": [\\\n    \"<string>\"\\\n  ],\n  \"llm_extract\": {\n    \"schema\": {}\n  },\n  \"links_on_page\": {\n    \"query_to_order_links_by\": \"<string>\",\n    \"include_links\": [\\\n      \"<string>\"\\\n    ],\n    \"exclude_links\": [\\\n      \"<string>\"\\\n    ]\n  },\n  \"screen_size\": {\n    \"screen_width\": 123,\n    \"screen_height\": 123\n  },\n  \"screenshot\": {\n    \"full_page\": true\n  },\n  \"metadata\": {},\n  \"max_age\": 0\n}\n'\n```\n\n```\nimport requests\n\nurl = \"https://api.olostep.com/v1/scrapes\"\n\npayload = {\n    \"url_to_scrape\": \"<string>\",\n    \"wait_before_scraping\": 123,\n    \"formats\": [],\n    \"actions\": [\\\n        {\\\n            \"type\": \"wait\",\\\n            \"milliseconds\": 1\\\n        }\\\n    ],\n    \"country\": \"<string>\",\n    \"remove_images\": False,\n    \"remove_class_names\": [\"<string>\"],\n    \"llm_extract\": { \"schema\": {} },\n    \"links_on_page\": {\n        \"query_to_order_links_by\": \"<string>\",\n        \"include_links\": [\"<string>\"],\n        \"exclude_links\": [\"<string>\"]\n    },\n    \"screen_size\": {\n        \"screen_width\": 123,\n        \"screen_height\": 123\n    },\n    \"screenshot\": { \"full_page\": True },\n    \"metadata\": {},\n    \"max_age\": 0\n}\nheaders = {\n    \"Authorization\": \"Bearer <token>\",\n    \"Content-Type\": \"application/json\"\n}\n\nresponse = requests.post(url, json=payload, headers=headers)\n\nprint(response.text)\n```\n\n```\nconst options = {\n  method: 'POST',\n  headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},\n  body: JSON.stringify({\n    url_to_scrape: '<string>',\n    wait_before_scraping: 123,\n    formats: [],\n    actions: [{type: 'wait', milliseconds: 1}],\n    country: '<string>',\n    remove_images: false,\n    remove_class_names: ['<string>'],\n    llm_extract: {schema: {}},\n    links_on_page: {\n      query_to_order_links_by: '<string>',\n      include_links: ['<string>'],\n      exclude_links: ['<string>']\n    },\n    screen_size: {screen_width: 123, screen_height: 123},\n    screenshot: {full_page: true},\n    metadata: {},\n    max_age: 0\n  })\n};\n\nfetch('https://api.olostep.com/v1/scrapes', options)\n  .then(res => res.json())\n  .then(res => console.log(res))\n  .catch(err => console.error(err));\n```\n\n```\n<?php\n\n$curl = curl_init();\n\ncurl_setopt_array($curl, [\\\n  CURLOPT_URL => \"https://api.olostep.com/v1/scrapes\",\\\n  CURLOPT_RETURNTRANSFER => true,\\\n  CURLOPT_ENCODING => \"\",\\\n  CURLOPT_MAXREDIRS => 10,\\\n  CURLOPT_TIMEOUT => 30,\\\n  CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,\\\n  CURLOPT_CUSTOMREQUEST => \"POST\",\\\n  CURLOPT_POSTFIELDS => json_encode([\\\n    'url_to_scrape' => '<string>',\\\n    'wait_before_scraping' => 123,\\\n    'formats' => [\\\n\\\n    ],\\\n    'actions' => [\\\n        [\\\n                'type' => 'wait',\\\n                'milliseconds' => 1\\\n        ]\\\n    ],\\\n    'country' => '<string>',\\\n    'remove_images' => false,\\\n    'remove_class_names' => [\\\n        '<string>'\\\n    ],\\\n    'llm_extract' => [\\\n        'schema' => [\\\n\\\n        ]\\\n    ],\\\n    'links_on_page' => [\\\n        'query_to_order_links_by' => '<string>',\\\n        'include_links' => [\\\n                '<string>'\\\n        ],\\\n        'exclude_links' => [\\\n                '<string>'\\\n        ]\\\n    ],\\\n    'screen_size' => [\\\n        'screen_width' => 123,\\\n        'screen_height' => 123\\\n    ],\\\n    'screenshot' => [\\\n        'full_page' => true\\\n    ],\\\n    'metadata' => [\\\n\\\n    ],\\\n    'max_age' => 0\\\n  ]),\\\n  CURLOPT_HTTPHEADER => [\\\n    \"Authorization: Bearer <token>\",\\\n    \"Content-Type: application/json\"\\\n  ],\\\n]);\n\n$response = curl_exec($curl);\n$err = curl_error($curl);\n\ncurl_close($curl);\n\nif ($err) {\n  echo \"cURL Error #:\" . $err;\n} else {\n  echo $response;\n}\n```\n\n```\npackage main\n\nimport (\n\t\"fmt\"\n\t\"strings\"\n\t\"net/http\"\n\t\"io\"\n)\n\nfunc main() {\n\n\turl := \"https://api.olostep.com/v1/scrapes\"\n\n\tpayload := strings.NewReader(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n\n\treq, _ := http.NewRequest(\"POST\", url, payload)\n\n\treq.Header.Add(\"Authorization\", \"Bearer <token>\")\n\treq.Header.Add(\"Content-Type\", \"application/json\")\n\n\tres, _ := http.DefaultClient.Do(req)\n\n\tdefer res.Body.Close()\n\tbody, _ := io.ReadAll(res.Body)\n\n\tfmt.Println(string(body))\n\n}\n```\n\n```\nHttpResponse<String> response = Unirest.post(\"https://api.olostep.com/v1/scrapes\")\n  .header(\"Authorization\", \"Bearer <token>\")\n  .header(\"Content-Type\", \"application/json\")\n  .body(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n  .asString();\n```\n\n```\nrequire 'uri'\nrequire 'net/http'\n\nurl = URI(\"https://api.olostep.com/v1/scrapes\")\n\nhttp = Net::HTTP.new(url.host, url.port)\nhttp.use_ssl = true\n\nrequest = Net::HTTP::Post.new(url)\nrequest[\"Authorization\"] = 'Bearer <token>'\nrequest[\"Content-Type\"] = 'application/json'\nrequest.body = \"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\"\n\nresponse = http.request(request)\nputs response.read_body\n```\n\n200\n\n400\n\n502\n\n504\n\n```\n{\n  \"id\": \"<string>\",\n  \"object\": \"<string>\",\n  \"created\": 123,\n  \"metadata\": {},\n  \"url_to_scrape\": \"<string>\",\n  \"result\": {\n    \"html_content\": \"<string>\",\n    \"markdown_content\": \"<string>\",\n    \"text_content\": \"<string>\",\n    \"json_content\": \"<string>\",\n    \"screenshot_hosted_url\": \"<string>\",\n    \"html_hosted_url\": \"<string>\",\n    \"markdown_hosted_url\": \"<string>\",\n    \"text_hosted_url\": \"<string>\",\n    \"links_on_page\": [\\\n      \"<string>\"\\\n    ],\n    \"page_metadata\": {\n      \"status_code\": 123,\n      \"title\": \"<string>\"\n    }\n  },\n  \"credits_consumed\": 123,\n  \"cost_usd\": 123\n}\n```\n\n```\n{\n  \"id\": \"error_x2nmu5bqn6\",\n  \"object\": \"error\",\n  \"created\": 1777923912,\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"dns_resolution_failed\",\n    \"message\": \"The URL contains a typo, or the domain does not exist.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_ogeb6rik8c\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"tls_error\",\n    \"detail\": \"err_ssl_tlsv1_alert_internal_error\",\n    \"message\": \"The website closed or rejected the TLS handshake. The server may be misconfigured or use an unsupported SSL/TLS version.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_qat3d1amjt\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"request_timeout\",\n    \"code\": \"scrape_poll_timeout\",\n    \"message\": \"Request timed out while waiting for scrape result. The page may be slow, blocked for our fetchers, or temporarily unavailable.\"\n  }\n}\n```\n\nPOST\n\n/\n\nv1\n\n/\n\nscrapes\n\nTry it\n\nInitiate a web page scrape\n\ncURL\n\n```\ncurl --request POST \\\n  --url https://api.olostep.com/v1/scrapes \\\n  --header 'Authorization: Bearer <token>' \\\n  --header 'Content-Type: application/json' \\\n  --data '\n{\n  \"url_to_scrape\": \"<string>\",\n  \"wait_before_scraping\": 123,\n  \"formats\": [],\n  \"actions\": [\\\n    {\\\n      \"type\": \"wait\",\\\n      \"milliseconds\": 1\\\n    }\\\n  ],\n  \"country\": \"<string>\",\n  \"remove_images\": false,\n  \"remove_class_names\": [\\\n    \"<string>\"\\\n  ],\n  \"llm_extract\": {\n    \"schema\": {}\n  },\n  \"links_on_page\": {\n    \"query_to_order_links_by\": \"<string>\",\n    \"include_links\": [\\\n      \"<string>\"\\\n    ],\n    \"exclude_links\": [\\\n      \"<string>\"\\\n    ]\n  },\n  \"screen_size\": {\n    \"screen_width\": 123,\n    \"screen_height\": 123\n  },\n  \"screenshot\": {\n    \"full_page\": true\n  },\n  \"metadata\": {},\n  \"max_age\": 0\n}\n'\n```\n\n```\nimport requests\n\nurl = \"https://api.olostep.com/v1/scrapes\"\n\npayload = {\n    \"url_to_scrape\": \"<string>\",\n    \"wait_before_scraping\": 123,\n    \"formats\": [],\n    \"actions\": [\\\n        {\\\n            \"type\": \"wait\",\\\n            \"milliseconds\": 1\\\n        }\\\n    ],\n    \"country\": \"<string>\",\n    \"remove_images\": False,\n    \"remove_class_names\": [\"<string>\"],\n    \"llm_extract\": { \"schema\": {} },\n    \"links_on_page\": {\n        \"query_to_order_links_by\": \"<string>\",\n        \"include_links\": [\"<string>\"],\n        \"exclude_links\": [\"<string>\"]\n    },\n    \"screen_size\": {\n        \"screen_width\": 123,\n        \"screen_height\": 123\n    },\n    \"screenshot\": { \"full_page\": True },\n    \"metadata\": {},\n    \"max_age\": 0\n}\nheaders = {\n    \"Authorization\": \"Bearer <token>\",\n    \"Content-Type\": \"application/json\"\n}\n\nresponse = requests.post(url, json=payload, headers=headers)\n\nprint(response.text)\n```\n\n```\nconst options = {\n  method: 'POST',\n  headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},\n  body: JSON.stringify({\n    url_to_scrape: '<string>',\n    wait_before_scraping: 123,\n    formats: [],\n    actions: [{type: 'wait', milliseconds: 1}],\n    country: '<string>',\n    remove_images: false,\n    remove_class_names: ['<string>'],\n    llm_extract: {schema: {}},\n    links_on_page: {\n      query_to_order_links_by: '<string>',\n      include_links: ['<string>'],\n      exclude_links: ['<string>']\n    },\n    screen_size: {screen_width: 123, screen_height: 123},\n    screenshot: {full_page: true},\n    metadata: {},\n    max_age: 0\n  })\n};\n\nfetch('https://api.olostep.com/v1/scrapes', options)\n  .then(res => res.json())\n  .then(res => console.log(res))\n  .catch(err => console.error(err));\n```\n\n```\n<?php\n\n$curl = curl_init();\n\ncurl_setopt_array($curl, [\\\n  CURLOPT_URL => \"https://api.olostep.com/v1/scrapes\",\\\n  CURLOPT_RETURNTRANSFER => true,\\\n  CURLOPT_ENCODING => \"\",\\\n  CURLOPT_MAXREDIRS => 10,\\\n  CURLOPT_TIMEOUT => 30,\\\n  CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,\\\n  CURLOPT_CUSTOMREQUEST => \"POST\",\\\n  CURLOPT_POSTFIELDS => json_encode([\\\n    'url_to_scrape' => '<string>',\\\n    'wait_before_scraping' => 123,\\\n    'formats' => [\\\n\\\n    ],\\\n    'actions' => [\\\n        [\\\n                'type' => 'wait',\\\n                'milliseconds' => 1\\\n        ]\\\n    ],\\\n    'country' => '<string>',\\\n    'remove_images' => false,\\\n    'remove_class_names' => [\\\n        '<string>'\\\n    ],\\\n    'llm_extract' => [\\\n        'schema' => [\\\n\\\n        ]\\\n    ],\\\n    'links_on_page' => [\\\n        'query_to_order_links_by' => '<string>',\\\n        'include_links' => [\\\n                '<string>'\\\n        ],\\\n        'exclude_links' => [\\\n                '<string>'\\\n        ]\\\n    ],\\\n    'screen_size' => [\\\n        'screen_width' => 123,\\\n        'screen_height' => 123\\\n    ],\\\n    'screenshot' => [\\\n        'full_page' => true\\\n    ],\\\n    'metadata' => [\\\n\\\n    ],\\\n    'max_age' => 0\\\n  ]),\\\n  CURLOPT_HTTPHEADER => [\\\n    \"Authorization: Bearer <token>\",\\\n    \"Content-Type: application/json\"\\\n  ],\\\n]);\n\n$response = curl_exec($curl);\n$err = curl_error($curl);\n\ncurl_close($curl);\n\nif ($err) {\n  echo \"cURL Error #:\" . $err;\n} else {\n  echo $response;\n}\n```\n\n```\npackage main\n\nimport (\n\t\"fmt\"\n\t\"strings\"\n\t\"net/http\"\n\t\"io\"\n)\n\nfunc main() {\n\n\turl := \"https://api.olostep.com/v1/scrapes\"\n\n\tpayload := strings.NewReader(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n\n\treq, _ := http.NewRequest(\"POST\", url, payload)\n\n\treq.Header.Add(\"Authorization\", \"Bearer <token>\")\n\treq.Header.Add(\"Content-Type\", \"application/json\")\n\n\tres, _ := http.DefaultClient.Do(req)\n\n\tdefer res.Body.Close()\n\tbody, _ := io.ReadAll(res.Body)\n\n\tfmt.Println(string(body))\n\n}\n```\n\n```\nHttpResponse<String> response = Unirest.post(\"https://api.olostep.com/v1/scrapes\")\n  .header(\"Authorization\", \"Bearer <token>\")\n  .header(\"Content-Type\", \"application/json\")\n  .body(\"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\")\n  .asString();\n```\n\n```\nrequire 'uri'\nrequire 'net/http'\n\nurl = URI(\"https://api.olostep.com/v1/scrapes\")\n\nhttp = Net::HTTP.new(url.host, url.port)\nhttp.use_ssl = true\n\nrequest = Net::HTTP::Post.new(url)\nrequest[\"Authorization\"] = 'Bearer <token>'\nrequest[\"Content-Type\"] = 'application/json'\nrequest.body = \"{\\n  \\\"url_to_scrape\\\": \\\"<string>\\\",\\n  \\\"wait_before_scraping\\\": 123,\\n  \\\"formats\\\": [],\\n  \\\"actions\\\": [\\n    {\\n      \\\"type\\\": \\\"wait\\\",\\n      \\\"milliseconds\\\": 1\\n    }\\n  ],\\n  \\\"country\\\": \\\"<string>\\\",\\n  \\\"remove_images\\\": false,\\n  \\\"remove_class_names\\\": [\\n    \\\"<string>\\\"\\n  ],\\n  \\\"llm_extract\\\": {\\n    \\\"schema\\\": {}\\n  },\\n  \\\"links_on_page\\\": {\\n    \\\"query_to_order_links_by\\\": \\\"<string>\\\",\\n    \\\"include_links\\\": [\\n      \\\"<string>\\\"\\n    ],\\n    \\\"exclude_links\\\": [\\n      \\\"<string>\\\"\\n    ]\\n  },\\n  \\\"screen_size\\\": {\\n    \\\"screen_width\\\": 123,\\n    \\\"screen_height\\\": 123\\n  },\\n  \\\"screenshot\\\": {\\n    \\\"full_page\\\": true\\n  },\\n  \\\"metadata\\\": {},\\n  \\\"max_age\\\": 0\\n}\"\n\nresponse = http.request(request)\nputs response.read_body\n```\n\n200\n\n400\n\n502\n\n504\n\n```\n{\n  \"id\": \"<string>\",\n  \"object\": \"<string>\",\n  \"created\": 123,\n  \"metadata\": {},\n  \"url_to_scrape\": \"<string>\",\n  \"result\": {\n    \"html_content\": \"<string>\",\n    \"markdown_content\": \"<string>\",\n    \"text_content\": \"<string>\",\n    \"json_content\": \"<string>\",\n    \"screenshot_hosted_url\": \"<string>\",\n    \"html_hosted_url\": \"<string>\",\n    \"markdown_hosted_url\": \"<string>\",\n    \"text_hosted_url\": \"<string>\",\n    \"links_on_page\": [\\\n      \"<string>\"\\\n    ],\n    \"page_metadata\": {\n      \"status_code\": 123,\n      \"title\": \"<string>\"\n    }\n  },\n  \"credits_consumed\": 123,\n  \"cost_usd\": 123\n}\n```\n\n```\n{\n  \"id\": \"error_x2nmu5bqn6\",\n  \"object\": \"error\",\n  \"created\": 1777923912,\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"dns_resolution_failed\",\n    \"message\": \"The URL contains a typo, or the domain does not exist.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_ogeb6rik8c\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"invalid_request_error\",\n    \"code\": \"tls_error\",\n    \"detail\": \"err_ssl_tlsv1_alert_internal_error\",\n    \"message\": \"The website closed or rejected the TLS handshake. The server may be misconfigured or use an unsupported SSL/TLS version.\"\n  }\n}\n```\n\n```\n{\n  \"id\": \"error_qat3d1amjt\",\n  \"object\": \"error\",\n  \"created\": 1777923969,\n  \"url\": \"https://example.com\",\n  \"metadata\": {},\n  \"error\": {\n    \"type\": \"request_timeout\",\n    \"code\": \"scrape_poll_timeout\",\n    \"message\": \"Request timed out while waiting for scrape result. The page may be slow, blocked for our fetchers, or temporarily unavailable.\"\n  }\n}\n```\n\n**Optional caching:** Pass `max_age` (in seconds) to reuse a recent scrape with the same parameters instead of fetching the page again. Defaults to `0` (always fresh). In the dashboard playground, the default is 24 hours. See [Caching](https://docs.olostep.com/features/scrapes#caching) for details.\n\n#### Authorizations\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#authorization-authorization)\n\nAuthorization\n\nstring\n\nheader\n\nrequired\n\nBearer authentication header of the form Bearer , where  is your auth token.\n\n#### Body\n\napplication/json\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-url-to-scrape)\n\nurl\\_to\\_scrape\n\nstring<uri>\n\nrequired\n\nThe URL to start scraping from.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-wait-before-scraping)\n\nwait\\_before\\_scraping\n\ninteger\n\nTime to wait in milliseconds before starting the scraping.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-formats)\n\nformats\n\nenum<string>\\[\\]\n\nFormats in which you want the content.\n\nAvailable options:\n\n`html`,\n\n`markdown`,\n\n`text`,\n\n`json`,\n\n`raw_pdf`,\n\n`screenshot`\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-remove-css-selectors)\n\nremove\\_css\\_selectors\n\nenum<string>\n\nOption to remove certain CSS selectors from the content. Optionally, you can also pass a JSON stringified array of specific selectors you want to remove. The CSS selectors removed when this option is set to default are \\['nav','footer','script','style','noscript','svg',\\[role=alert\\],\\[role=banner\\],\\[role=dialog\\],\\[role=alertdialog\\],\\[role=region\\]\\[aria-label\\*=skip i\\],\\[aria-modal=true\\]\\]\n\nAvailable options:\n\n`default`,\n\n`none`,\n\n`array`\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-actions)\n\nactions\n\n(Wait \u00b7 object \\| Click \u00b7 object \\| Fill Input \u00b7 object \\| Scroll \u00b7 object)\\[\\]\n\nActions to perform on the page before getting the content.\n\n- Wait\n\n- Click\n\n- Fill Input\n\n- Scroll\n\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-country)\n\ncountry\n\nstring\n\nResidential country to load the request from.\n\nSupported values are:\n\n- US (United States)\n- CA (Canada)\n- IT (Italy)\n- IN (India)\n- GB (England)\n- JP (Japan)\n- MX (Mexico)\n- AU (Australia)\n- ID (Indonesia)\n- UA (UAE)\n- RU (Russia)\n- RANDOM\n\nSome operations, like scraping Google Search and Google News, support all countries.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-transformer)\n\ntransformer\n\nenum<string>\n\nSpecify the HTML transformer to use, if any. Postlight's Mercury Parser library is used to remove ads and other unwanted content from the scraped content.\n\nAvailable options:\n\n`postlight`,\n\n`none`\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-remove-images)\n\nremove\\_images\n\nboolean\n\ndefault:false\n\nOption to remove images from the scraped content. Defaults to false.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-remove-class-names)\n\nremove\\_class\\_names\n\nstring\\[\\]\n\nList of class names to remove from the content.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-parser)\n\nparser\n\nobject\n\nWhen defining json as a format, you can use this parameter to specify the parser to use. Parsers are useful to extract structured content from web pages. Olostep has a few parsers built in for most common web pages, and you can also create your own parsers.\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-llm-extract)\n\nllm\\_extract\n\nobject\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-links-on-page)\n\nlinks\\_on\\_page\n\nobject\n\nWith this option, you can get all the links present on the page you scrape. Links are always returned as absolute URLs.\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-screen-size)\n\nscreen\\_size\n\nobject\n\nConfiguration for screen size. Preset dimensions are available through screen\\_type: desktop (1920x1080), mobile (414x896), or default (768x1024).\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-screenshot)\n\nscreenshot\n\nobject\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-metadata)\n\nmetadata\n\nobject\n\nUser-defined metadata. Not supported yet\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#body-max-age)\n\nmax\\_age\n\ninteger\n\ndefault:0\n\nMaximum acceptable age of cached content, in seconds. When a matching scrape already exists and is newer than max\\_age seconds, Olostep returns the stored result instead of launching a new browser scrape. Defaults to 0 (always scrape fresh). In the dashboard playground, the default is 86400 (24 hours). The maximum allowed value is 604800 (7 days). See the Caching section in the Scrapes feature docs for details.\n\nRequired range: `x >= 0`\n\n#### Response\n\n200\n\napplication/json\n\nSuccessful response with the scrape initiation details.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-id)\n\nid\n\nstring\n\nScrape ID\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-object)\n\nobject\n\nstring\n\nThe kind of object. \"scrape\" for this endpoint.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-created)\n\ncreated\n\nnumber\n\nCreated epoch\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-metadata)\n\nmetadata\n\nobject\n\nUser-defined metadata.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-url-to-scrape)\n\nurl\\_to\\_scrape\n\nstring\n\nThe URL that was scraped.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-result)\n\nresult\n\nobject\n\nShowchild attributes\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-credits-consumed-one-of-0)\n\ncredits\\_consumed\n\ninteger \\| null\n\nNumber of credits consumed by this request. Populated after execution completes. Credits are the source of truth for billing.\n\n[\u200b](https://docs.olostep.com/api-reference/scrapes/create#response-cost-usd-one-of-0)\n\ncost\\_usd\n\nnumber \\| null\n\nEstimated cost in USD for this request. Populated after execution completes. Calculated from credits consumed and your plan rate \u2014 99% accurate, but credits\\_consumed is the authoritative value.\n\nWas this page helpful?\n\nYesNo",
          "metadata": {
            "language": "en",
            "generator": "Mintlify",
            "og:description": "Scrape a url with provided configuration and get content.",
            "og:site_name": "Olostep Docs",
            "msapplication-config": "/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/browserconfig.xml",
            "og:type": "website",
            "og:image:width": "1200",
            "twitter:card": "summary_large_image",
            "ogImage": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DScrapes%26title%3DCreate%2BScrape%26description%3D%255BScrape%255D%2528https%253A%252F%252Fdocs.olostep.com%252Ffeatures%252Fscrapes%2529%2Ba%2Burl%2Bwith%2Bprovided%2Bconfiguration%2Band%2Bget%2Bcontent.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "robots": "index, follow",
            "ogUrl": "https://docs.olostep.com/api-reference/scrapes/create",
            "twitter:image:height": "630",
            "description": "Scrape a url with provided configuration and get content.",
            "title": "Create Scrape - Olostep Docs",
            "og:title": "Create Scrape - Olostep Docs",
            "ogDescription": "Scrape a url with provided configuration and get content.",
            "twitter:image:width": "1200",
            "ogSiteName": "Olostep Docs",
            "og:image:height": "630",
            "twitter:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DScrapes%26title%3DCreate%2BScrape%26description%3D%255BScrape%255D%2528https%253A%252F%252Fdocs.olostep.com%252Ffeatures%252Fscrapes%2529%2Ba%2Burl%2Bwith%2Bprovided%2Bconfiguration%2Band%2Bget%2Bcontent.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "apple-mobile-web-app-title": "Olostep Docs",
            "viewport": "width=device-width, initial-scale=1, viewport-fit=cover",
            "application-name": "Olostep Docs",
            "twitter:title": "Create Scrape - Olostep Docs",
            "ogTitle": "Create Scrape - Olostep Docs",
            "twitter:description": "Scrape a url with provided configuration and get content.",
            "og:url": "https://docs.olostep.com/api-reference/scrapes/create",
            "og:image": "https://olostep-58.mintlify.app/mintlify-assets/_next/image?url=%2F_mintlify%2Fapi%2Fog%3Fdivision%3DScrapes%26title%3DCreate%2BScrape%26description%3D%255BScrape%255D%2528https%253A%252F%252Fdocs.olostep.com%252Ffeatures%252Fscrapes%2529%2Ba%2Burl%2Bwith%2Bprovided%2Bconfiguration%2Band%2Bget%2Bcontent.%26theme%3Ddf586dd123820a765d726092&w=1200&q=100",
            "msapplication-TileColor": "#9563FF",
            "googlebot": "index, follow",
            "favicon": "https://docs.olostep.com/mintlify-assets/_mintlify/favicons/olostep-58/_83ELNnJOV2rn47b/_generated/favicon/android-chrome-192x192.png",
            "scrapeId": "01a06bbe-7eb7-745a-98ef-b55c3432b37b",
            "sourceURL": "https://docs.olostep.com/api-reference/scrapes/create",
            "url": "https://docs.olostep.com/api-reference/scrapes/create",
            "statusCode": 200,
            "contentType": "text/html; charset=utf-8",
            "timezone": "America/New_York",
            "proxyUsed": "basic",
            "cacheState": "miss",
            "indexId": "951b7e3d-ee17-428e-baba-8731eb83fa72",
            "creditsUsed": 1
          }
        }
      ]
    },
    "creditsUsed": 12,
    "id": "01a06bbe-7d66-712f-98a6-f7f88f0a9b51"
  }
}