API Documentation
Endpoint
POST
https://urltotext.com/api/v1/urltotext/
Headers
- Authorization Token <API_TOKEN>
- Content-Type application/json
Request Payload
| Parameter | Type | Required | Description |
|---|---|---|---|
| url | string | true | The URL of the webpage or YouTube video to convert. |
| output_format | string | false | The format to return the content in. Choices are "text" or "markdown" or "html". Default is "text". |
| extract_main_content | boolean | false | If true, attempts to extract only the main content using AI. Default is false. |
| render_javascript | boolean | false | If true, renders JavaScript on the page before extracting content. Default is true. |
| residential_proxy | boolean | false | If true, uses a residential proxy for the request. Note: Only one proxy option can be selected at a time (residential_proxy or stealth_proxy). If both are selected, stealth_proxy takes priority. Default is false. |
| ai_prompt | string | false | An optional AI prompt to process or modify the extracted content. |
| end_of_article | string | false | Optional string marking where to stop extracting content in the article. |
| wait_for_js | integer | false | Specifies the duration (in milliseconds) to wait for JavaScript to fully load on the page before fetching content. This is particularly useful for pages with dynamic content that may take time to render. |
| stealth_proxy | boolean | false | If true, uses our pool of premium residential IP addresses to bypass very difficult-to-scrape websites. Note: Only one proxy option can be selected at a time (residential_proxy or stealth_proxy). If both are selected, stealth_proxy takes priority. Default is false. It requires render_javascript to be set as true. |
| extract_css_selector | string | false | Enter one or more valid CSS selectors separated by commas (e.g., '#main-content, .accordion'). These selectors will be used to extract only the specified parts of the webpage. If left empty or not matching selectors found, the entire page content will be processed. |
Response Codes
- 200: Success
- 400: Bad Request (Invalid parameters or request body)
- 402: Payment Required (Insufficient credits)
- 422: Invalid Request (Field Errors)
- 500: Internal Server Error (Server failure)
Code Examples
curl --location 'https://urltotext.com/api/v1/urltotext/' \
--header 'Authorization: Token YOUR_API_TOKEN' \
--header 'Content-Type: application/json' \
--data '{
"url": "https://example.com",
"output_format": "markdown",
"extract_main_content": true,
"render_javascript": true,
"residential_proxy": false
}'
import requests
import json
url = "https://urltotext.com/api/v1/urltotext/"
payload = json.dumps(
{
"url": "https://example.com",
"output_format": "markdown",
"extract_main_content": True,
"render_javascript": True,
"residential_proxy": False,
}
)
headers = {
"Authorization": "Token YOUR_API_TOKEN",
"Content-Type": "application/json"
}
response = requests.request("POST", url, headers=headers, data=payload)
print(response.text)
fetch('https://urltotext.com/api/v1/urltotext/', {
method: 'POST',
headers: {
'Authorization': 'Token YOUR_API_TOKEN',
'Content-Type': 'application/json'
},
body: JSON.stringify({
url: 'https://example.com',
output_format: 'markdown',
extract_main_content: true,
render_javascript: true,
residential_proxy: false
})
})
.then(response => response.json())
.then(data => console.log(data))
.catch(error => console.error(error));
<?php
$data = [
'url' => 'https://example.com',
'output_format' => 'markdown',
'extract_main_content' => true,
'render_javascript' => true,
'residential_proxy' => false
];
$ch = curl_init('https://urltotext.com/api/v1/urltotext/');
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => json_encode($data),
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
'Authorization: Token YOUR_API_TOKEN',
'Content-Type: application/json'
]
]);
$response = curl_exec($ch);
curl_close($ch);
echo $response;
Sample Response
{
"message": "Request processed successfully",
"data": {
"url": "https://example.com",
"page_title": "title of the page",
"published_date": "2024-12-05T10:15:00",
"og_image_url": "open grapgh image url",
"og_description": "open graph description",
"output_format": "markdown",
"content": "# Sample Content"
"warning:" " # Sample Warning"
},
"credits_used": "0.02"
}
Endpoint
POST
https://urltotext.com/pdf_to_text/api/v1/convert/
Headers
- Authorization Token <API_TOKEN>
- Content-Type multipart/form-data
Request Payload
| Parameter | Type | Required | Description |
|---|---|---|---|
| file | file | true | The PDF file to convert to text. Maximum file size is 10MB. |
Response Codes
- 200: Success
- 400: Bad Request (Invalid file or file too large)
- 401: Unauthorized (Invalid or missing API token)
- 402: Payment Required (Insufficient credits)
- 500: Internal Server Error (Server failure)
Code Examples
curl --location 'https://urltotext.com/pdf_to_text/api/v1/convert/' \ --header 'Authorization: Token YOUR_API_TOKEN' \ --form 'file=@"/path/to/your/file.pdf"'
import requests
url = "https://urltotext.com/pdf_to_text/api/v1/convert/"
files = {
'file': open('/path/to/your/file.pdf', 'rb')
}
headers = {
"Authorization": "Token YOUR_API_TOKEN"
}
response = requests.post(url, headers=headers, files=files)
print(response.text)
const formData = new FormData();
formData.append('file', fileInput.files[0]);
fetch('https://urltotext.com/pdf_to_text/api/v1/convert/', {
method: 'POST',
headers: {
'Authorization': 'Token YOUR_API_TOKEN'
},
body: formData
})
.then(response => response.json())
.then(data => console.log(data))
.catch(error => console.error(error));
<?php
$file = '/path/to/your/file.pdf';
$ch = curl_init('https://urltotext.com/pdf_to_text/api/v1/convert/');
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => [
'file' => new CURLFile($file, 'application/pdf', basename($file))
],
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
'Authorization: Token YOUR_API_TOKEN'
]
]);
$response = curl_exec($ch);
curl_close($ch);
echo $response;
Sample Response
{
"message": "PDF converted successfully",
"data": {
"text": "This is the extracted text content from the PDF file...",
"warning": "Some text may not have been extracted properly due to image-based content"
},
"credits_used": "0.01"
}
Learn More about URL to Text
Welcome to URL to Text! We're excited to introduce you to our platform through this brief video.
We know this is annoying, but getting this information is very important for us to help you better. Thank you for taking the time to share how you discovered us.
This field is required.