N
News Els — Web Crawler Architecture

Polite & Ethical Web Crawling Architecture

A comprehensive blueprint for automated global news aggregation. Designed to crawl respectful of target site resources, comply strictly with web standards, protect site owners' interests, avoid IP blacklisting, and collate daily bite-sized news digests for users.

📑 Document Contents

1. Mission & Core Philosophy

2. The 6 Pillars of Polite Crawling

3. System Architecture & Data Flow

4. RSS / Atom Feed Preference Strategy

5. Concurrency & Rate Limiting Engine

6. Avoiding Blacklists & IP Bans

7. AI News Digest & Snippet Generation

8. Copyright, Attribution & Fair Use

1. Mission & Core Philosophy

The goal of the News Els Engine is to gather, collate, and deliver fresh, daily bite-sized news digests customized to user preferences across categories (Tech, World, Finance, Science, Sports).

💡 Core Ethical Principle: Our crawlers must act as exemplary web citizens. We prioritize target server stability, strictly observe bandwidth boundaries, respect publisher rights, and never degrade the performance or experience of original web hosts.

2. The 6 Pillars of Polite Crawling

1 robots.txt Compliance

Always fetch, parse, and strictly obey robots.txt directives (including Disallow, Allow, and Crawl-delay) prior to requesting any path on a domain.

2 Domain Rate Limiting

Enforce strict per-host token bucket queues. Limit requests to max 1 request every 2 to 5 seconds per domain and cap concurrent TCP sockets per target host to 1.

3 Transparent Bot Identity

Declare explicit User-Agent strings containing our bot name, version, project URL, and contact email so publishers can easily identify or contact us.

4 Bandwidth Efficiency

Prefer RSS/Atom XML feeds over HTML scraping. Use HTTP ETag & If-Modified-Since headers to receive 304 Not Modified responses and skip re-downloading unchanged content.

5 Circuit Breaker & Backoff

Detect HTTP 429 (Too Many Requests) or 503 (Service Unavailable). Immediately trigger exponential backoff with random jitter and cool down the host for 30–60 minutes.

6 Fair Use & Attribution

Never mirror full article bodies or steal original content. Store only canonical URLs, publication dates, and generate short 2-sentence AI summaries linking back to the original source.

3. System Architecture & Data Flow

The crawler pipeline isolates queue scheduling, rate-limiting, network fetching, content parsing, and AI summarization into decoupled micro-stages.

1. Source Registry RSS Feeds & Web URLs 2. Polite Gatekeeper robots.txt + Crawl-Delay Per-Domain Token Bucket 3. Fetcher Engine ETag / Conditional GET 4. Content Extractor RSS XML / HTML Article 5. Deduplication URL Hash & Title Similarity 6. AI Snippet Generator 2-Sentence Bullet Summary 7. News Els Feed API User Category Digest

4. RSS / Atom Feed Preference Strategy

Scraping raw HTML pages is computationally heavy for target servers and fragile for crawlers. The News Els engine operates on a Feed-First Architecture:

Ingestion Method Server Overhead Bandwidth Used Politeness Rating
RSS 2.0 / Atom XML Feeds Minimal (Static cached XML) < 15 KB per fetch Optimal 🟢
Conditional HTTP GET (ETag / 304) Zero processing (Header check) < 1 KB (Headers only) Optimal 🟢
Selective HTML Scraping Moderate (HTML Rendering) ~100–500 KB per page Polite Controlled 🟡
Unthrottled Recursive Crawling High (Server stress / Spike) High bandwidth waste Strictly Forbidden 🔴

5. Concurrency & Rate Limiting Engine

To prevent any possibility of triggering a Denial-of-Service (DoS) condition on news publishers, our rate-limiter implements a Per-Domain Token Bucket pattern in memory:

// Token Bucket Rate Limiter Pseudocode for Domain Ingestion
class DomainRateLimiter {
  constructor() {
    this.domainQueues = new Map(); // domain -> Queue
    this.minDelayMs = 3000;         // Minimum 3 seconds between requests per domain
  }

  async fetchPolitely(targetUrl) {
    const domain = new URL(targetUrl).hostname;
    await this.acquireDomainLock(domain);

    try {
      // 1. Check robots.txt compliance
      const robotsAllowed = await checkRobotsTxt(domain, targetUrl);
      if (!robotsAllowed) {
        console.warn(`[Robots.txt] Disallowed: ${targetUrl}`);
        return null;
      }

      // 2. Perform conditional HTTP request with User-Agent & ETag
      const response = await httpGetWithHeaders(targetUrl, {
        'User-Agent': 'NewsElsBot/1.0 (+https://news.elsastry.com/bot; bot@news.elsastry.com)',
        'If-None-Match': getCachedETag(targetUrl),
        'If-Modified-Since': getCachedLastModified(targetUrl)
      });

      return response;
    } finally {
      // Release lock with mandatory 3-second sleep before next request to same domain
      setTimeout(() => this.releaseDomainLock(domain), this.minDelayMs);
    }
  }
}

6. Avoiding Blacklists & IP Bans

IP bans and WAF blocks (Cloudflare, Akamai, Imperva) happen when crawlers exhibit aggressive, un-humanlike behavior. We employ these safeguard rules:

🛡️ Anti-Blacklisting Checklist

7. AI News Digest & Snippet Generation

Users want clean, fresh, daily updates without reading long wall-of-text articles. The News Els Summarizer processes raw articles into 2-sentence key digests:

📰 Sample Snippet Card Output

TECHNOLOGY Source: TechCrunch • 15 mins ago

Quantum Computing Breakthrough Achieves Room-Temperature Coherence

Researchers have demonstrated a novel synthetic diamond matrix capable of maintaining qubit coherence at ambient temperatures for over 5 seconds. This breakthrough could drastically lower the cooling costs required for commercial quantum processors.

⏱️ 1 Min Read • 2-Sentence AI Summary Read Full Story on TechCrunch →

🔗 Direct Source Attribution

Every news snippet prominently displays the publisher's logo, brand name, and direct canonical link to the original article page.

📝 Original AI Summaries

We never copy-paste full article text. We generate concise, original 2-sentence bullet points summarizing key facts.

🚫 No Content Gating Bypass

Our crawlers never bypass paywalls, subscription gates, or login screens. If content is gated, it is excluded from ingestion.

📬 Publisher Opt-Out Process

Any site administrator can email bot-optout@news.elsastry.com or add User-agent: NewsElsBot Disallow: / to immediately stop ingestion.