Робот TechnicalReferenceBot
Это страница нашего поискового робота. Если вы увидели его в журнале своего сайта и хотите понять, кто пришёл и как его ограничить, — здесь всё написано.
User-agent: TechnicalReferenceBot/1.0 (+https://altkatalog.ru/bot)
Languages: English · Deutsch · Français · Español · 中文
Робот не маскируется под браузер и не выдаёт себя за чужого робота. Если в подписи стоит эта строка — это мы, и с нами можно связаться.
Что он собирает
Технические характеристики сложных товаров: компьютерных комплектующих, серверного и сетевого оборудования, бытовой техники, электроники и приборов. Собранное идёт в справочник, где у каждой характеристики указана страница, с которой она взята.
Берутся факты: производитель, модель, артикул, штрихкод, параметры и единицы их измерения, адрес страницы-источника. Чужие описания и фотографии оптом не забираются, персональные данные не собираются, закрытые разделы и личные кабинеты не трогаются.
Правила обхода
- robots.txt соблюдается. Запрещённые в нём разделы не запрашиваются вовсе.
- Crawl-delay соблюдается. Если он не задан — не чаще одного запроса в секунду к одному сайту.
- Один запрос за раз. К одному сайту робот обращается строго последовательно, сколько бы машин ни работало на нашей стороне.
- noindex, nofollow, meta robots, X-Robots-Tag, canonical учитываются.
- 429 и 503 останавливают обход всего сайта, а заголовок Retry-After соблюдается: попросили подождать — ждём.
- Повторно ничего не качается зря. Используются условные запросы по ETag и Last-Modified, а разбор идёт по уже сохранённой копии.
- Глубина и число страниц на сайт ограничены.
- Защита не обходится. Ни CAPTCHA, ни ограничения по частоте, ни любые технические запреты владельца ресурса.
Как ограничить или запретить обход
Обычным способом — файлом robots.txt в корне сайта. Полный запрет:
User-agent: TechnicalReferenceBot Disallow: /
Разрешить только каталог и попросить паузу побольше:
User-agent: TechnicalReferenceBot Allow: /catalog/ Disallow: / Crawl-delay: 10
Изменения вступают в силу в течение суток: столько robots.txt хранится у робота. Нужно быстрее — напишите, остановим руками.
Связь с оператором
Оператор робота — редакция «Альтернативы». Пишите на altkatalog@mail.ru, письма читает человек.
Напишите, если робот мешает работе сайта, берёт лишнее или, наоборот, вы хотите, чтобы данные обновлялись чаще. Просьбу об исключении выполняем без обсуждения условий и встречных предложений. Сообщите адрес страницы или товара — уберём из справочника и внесём в список исключений.
TechnicalReferenceBot — information for site owners
This is the crawler operated by Alternativa (altkatalog.ru). It identifies itself honestly and never pretends to be a browser or another crawler:
User-agent: TechnicalReferenceBot/1.0 (+https://altkatalog.ru/bot)
What it collects
Technical specifications of complex products: computer components, server and network hardware, home appliances, electronics and instruments. It gathers facts — manufacturer, model, part number, barcode, parameters and their units, and the address of the source page. Every specification in our reference keeps a link to the page it came from.
It does not harvest third-party descriptions or photographs in bulk, does not collect personal data, and does not touch private areas or user accounts.
Crawling rules
- robots.txt is obeyed. Disallowed sections are never requested.
- Crawl-delay is obeyed. Without one, at most one request per second per site.
- One request at a time per site, no matter how many machines we run.
- noindex, nofollow, meta robots, X-Robots-Tag and canonical are respected.
- 429 and 503 pause the whole site, and Retry-After is honoured.
- Nothing is re-downloaded needlessly: conditional requests with ETag and Last-Modified, and re-parsing happens on our stored copy.
- Crawl depth and page count per site are limited.
- No protection is circumvented — no CAPTCHA solving, no rate-limit evasion, no working around the owner's technical restrictions.
How to limit or block it
The ordinary way, via robots.txt in your site root. Block everything:
User-agent: TechnicalReferenceBot Disallow: /
Allow the catalogue only and ask for a longer pause:
User-agent: TechnicalReferenceBot Allow: /catalog/ Disallow: / Crawl-delay: 10
Changes take effect within 24 hours — that is how long the crawler keeps a cached robots.txt. If you need it sooner, write to us and we will stop it by hand.
Contact
Operator: the Alternativa team. Write to altkatalog@mail.ru — a human reads it.
Write if the crawler is a burden on your site, takes more than it should, or if you would rather your data were refreshed more often. Exclusion requests are honoured without negotiation. Send a page or product address and it will be removed from the reference and added to the exclusion list.
TechnicalReferenceBot — Hinweise für Website-Betreiber
Dies ist der Crawler von Alternativa (altkatalog.ru). Er gibt sich immer zu erkennen und tarnt sich weder als Browser noch als fremder Crawler:
User-agent: TechnicalReferenceBot/1.0 (+https://altkatalog.ru/bot)
Er sammelt technische Daten komplexer Produkte: Computerkomponenten, Server- und Netzwerktechnik, Haushaltsgeräte, Elektronik und Messtechnik. Erfasst werden Fakten — Hersteller, Modell, Artikelnummer, Barcode, Parameter mit Einheiten sowie die Adresse der Quellseite. Fremde Beschreibungen und Fotos werden nicht massenhaft übernommen, personenbezogene Daten nicht erhoben, geschlossene Bereiche nicht angetastet.
Regeln: robots.txt und Crawl-delay werden befolgt; ohne Angabe höchstens eine Anfrage pro Sekunde und Website; immer nur eine Anfrage gleichzeitig pro Website; noindex, nofollow, meta robots, X-Robots-Tag und canonical werden beachtet; 429 und 503 stoppen die gesamte Website, Retry-After wird eingehalten; Schutzmaßnahmen und CAPTCHAs werden nicht umgangen.
Sperren können Sie den Crawler wie gewohnt über robots.txt:
User-agent: TechnicalReferenceBot Disallow: /
Änderungen wirken innerhalb von 24 Stunden. Bei Problemen schreiben Sie an altkatalog@mail.ru — Ausschlusswünsche werden ohne Rückfragen umgesetzt.
TechnicalReferenceBot — informations pour les propriétaires de sites
Ce robot d'exploration est opéré par Alternativa (altkatalog.ru). Il s'identifie honnêtement et ne se fait jamais passer pour un navigateur ou pour un autre robot :
User-agent: TechnicalReferenceBot/1.0 (+https://altkatalog.ru/bot)
Il collecte les caractéristiques techniques de produits complexes : composants informatiques, matériel serveur et réseau, électroménager, électronique et instruments. Il recueille des faits — fabricant, modèle, référence, code-barres, paramètres et leurs unités, ainsi que l'adresse de la page source. Il ne récupère pas en masse les descriptions et photographies de tiers, ne collecte pas de données personnelles et ne touche pas aux espaces privés.
Règles : robots.txt et Crawl-delay sont respectés ; à défaut, une requête par seconde au maximum par site ; une seule requête à la fois par site ; noindex, nofollow, meta robots, X-Robots-Tag et canonical sont pris en compte ; 429 et 503 suspendent tout le site et Retry-After est respecté ; aucune protection ni CAPTCHA n'est contournée.
Pour le bloquer, utilisez robots.txt à la racine de votre site :
User-agent: TechnicalReferenceBot Disallow: /
Les modifications prennent effet sous 24 heures. En cas de problème, écrivez à altkatalog@mail.ru : les demandes d'exclusion sont satisfaites sans discussion.
TechnicalReferenceBot — información para propietarios de sitios
Este rastreador es operado por Alternativa (altkatalog.ru). Se identifica con honestidad y nunca se hace pasar por un navegador ni por otro rastreador:
User-agent: TechnicalReferenceBot/1.0 (+https://altkatalog.ru/bot)
Recopila características técnicas de productos complejos: componentes informáticos, equipos de servidor y de red, electrodomésticos, electrónica e instrumentos. Recoge hechos: fabricante, modelo, número de pieza, código de barras, parámetros con sus unidades y la dirección de la página de origen. No descarga masivamente descripciones ni fotografías ajenas, no recopila datos personales y no accede a áreas privadas.
Reglas: se respetan robots.txt y Crawl-delay; sin él, como máximo una petición por segundo y sitio; una sola petición simultánea por sitio; se respetan noindex, nofollow, meta robots, X-Robots-Tag y canonical; 429 y 503 detienen todo el sitio y se respeta Retry-After; no se elude ninguna protección ni CAPTCHA.
Para bloquearlo, use robots.txt en la raíz de su sitio:
User-agent: TechnicalReferenceBot Disallow: /
Los cambios surten efecto en 24 horas. Si algo va mal, escriba a altkatalog@mail.ru: las solicitudes de exclusión se atienden sin condiciones.
TechnicalReferenceBot — 网站管理员须知
本爬虫由 Alternativa(altkatalog.ru)运营。它如实标明身份,绝不伪装成浏览器或其他爬虫:
User-agent: TechnicalReferenceBot/1.0 (+https://altkatalog.ru/bot)
它收集复杂产品的技术参数:计算机配件、服务器与网络设备、家用电器、电子产品和仪器。 只采集事实数据——制造商、型号、料号、条码、参数及其单位,以及来源页面地址。不会批量抓取 他人的描述和图片,不收集个人数据,不访问受限区域和用户账户。
抓取规则:遵守 robots.txt 与 Crawl-delay;未设置时,对同一网站每秒最多一次请求; 对同一网站同时只发起一个请求;遵守 noindex、nofollow、meta robots、X-Robots-Tag 和 canonical;收到 429 或 503 时暂停整站抓取,并遵守 Retry-After;不绕过任何防护措施或验证码。
如需屏蔽,请在网站根目录的 robots.txt 中添加:
User-agent: TechnicalReferenceBot Disallow: /
更改将在 24 小时内生效。如有问题,请联系 altkatalog@mail.ru,我们会无条件执行排除请求。