### Development setup commands Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/CONTRIBUTING.md Clone the repository and install the package with development and optional dependencies. Use `.[dev]` for unit tests, `.[all]` for integration tests requiring document-extraction libraries and database drivers. ```bash git clone https://github.com/simonplmak-cloud/hkex-filing-scraper.git cd hkex-filing-scraper pip install -e ".[dev,all]" ``` -------------------------------- ### Install and publish to MCP registry Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/marketing/launch-v2.4.0.md One-time setup to install the mcp-publisher tool and publish the server to the MCP registry. Requires GitHub login. The server.json file is at the repo root and is schema-validated. After publishing, verify the listing with a curl command. ```bash # install curl -L "https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_$(uname -s | tr '[:upper:]' '[:lower:]')_$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/').tar.gz" | tar xz mcp-publisher && sudo mv mcp-publisher /usr/local/bin/ # publish (server.json is at the repo root and schema-validated) mcp-publisher login github mcp-publisher publish # verify curl -s "https://registry.modelcontextprotocol.io/v0/servers?search=hkex-filings" | head ``` -------------------------------- ### Install the release from PyPI Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Releasing Install the latest published version of hkex-filing-scraper from PyPI. ```bash pip install hkex-filing-scraper ``` -------------------------------- ### Install from PyPI or Git tag Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/docs/releasing.md Install the package from PyPI, or pin a specific release by installing directly from the Git tag. The second command installs the exact v1.2.0 tag from GitHub. ```bash pip install hkex-filing-scraper ``` ```bash pip install "git+https://github.com/simonplmak-cloud/hkex-filing-scraper@v1.2.0" ``` -------------------------------- ### Install from GitHub Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Getting-Started Installs the latest unreleased code directly from the GitHub repository. ```bash pip install "git+https://github.com/simonplmak-cloud/hkex-filing-scraper.git" ``` -------------------------------- ### Install core or all extras Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Getting-Started Installs the core package or the all-inclusive extra. The [all] extra includes Excel, dotenv, every database driver, and the MCP server. ```bash pip install hkex-filing-scraper # core; SQLite works out of the box pip install "hkex-filing-scraper[all]" # Excel + dotenv + every database driver + the MCP server ``` -------------------------------- ### Install and run hkex-scraper with metadata-only Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/README.md Install the package with pip, copy the environment template, and run a metadata-only scrape limited to 100 filings. The core install requires no server; the [all] extra adds Excel, dotenv, every driver, and the MCP server. ```bash pip install hkex-filing-scraper # core; SQLite needs no server pip install "hkex-filing-scraper[all]" # Excel + dotenv + every driver + the MCP server cp .env.example .env # then set DATABASE_TARGET (below) hkex-scraper --metadata-only --limit 100 ``` -------------------------------- ### Install MongoDB extra Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/MongoDB Install the MongoDB extra for the project. Required before using the sink. ```bash pip install ".[mongodb]" ``` -------------------------------- ### Install and run tests/lint for hkex-filing-scraper Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/AGENTS.md Installation and development commands for the hkex-filing-scraper package. The editable install with dev and all extras includes test/lint dependencies, doc extraction, and PostgreSQL/MySQL drivers; minimal install omits doc extraction and drivers but SQLite works. Optional drivers are added via extras like [postgres], [mysql], [duckdb], [mongodb], [clickhouse], or [neo4j]. Run pytest for tests (no DB/network needed) and ruff check for linting (target py310, line-length 100). ```bash pip install -e ".[dev,all]" # dev install (editable + test/lint deps + doc extraction + postgres + mysql) ``` ```bash pip install . # minimal install (no doc extraction, no drivers; SQLite works) ``` ```bash pip install ".[postgres]" # add the optional psycopg driver ``` ```bash pip install ".[mysql]" # add the optional PyMySQL driver (MySQL/MariaDB) ``` ```bash pip install ".[duckdb]" # or [mongodb], [clickhouse], [neo4j] ``` ```bash pytest # run all tests (no DB/network needed) ``` ```bash ruff check # lint (target: py310, line-length: 100) ``` -------------------------------- ### Install Neo4j sink extra Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Neo4j Install the optional neo4j dependency to enable the Neo4j sink. ```bash pip install ".[neo4j]" ``` -------------------------------- ### hkex-scraper CLI examples Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/CLI Provides common usage examples for backfilling documents, rebuilding graph edges, overriding database targets, dry-run validation, generating reports, verifying sinks, and checking version. ```bash # Backfill documents only, 50 at a time hkex-scraper --backfill-docs --limit 50 # Rebuild graph edges from existing filings hkex-scraper --link-only # Write to SQLite only for this run hkex-scraper --database-target postgres --limit 100 # Validate config without writing hkex-scraper --dry-run --limit 10 # Coverage and parity hkex-scraper --coverage-report hkex-scraper --database-target postgres,sqlite --parity-report # Cross-sink reconciliation (filing ids + document hashes) hkex-scraper --database-target postgres,sqlite --verify # Version hkex-scraper --version ``` -------------------------------- ### Run quickstart helper for a sink Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/examples/README.md Uses the helper script to write an `.env` for the named sink and run a metadata-only scrape. Supported sinks include postgres, sqlite, and mysql. ```bash ./examples/quickstart.sh postgres ./examples/quickstart.sh sqlite ./examples/quickstart.sh mysql ``` -------------------------------- ### Run MCP server Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/MCP Start the MCP server from the directory containing your .env and DATABASE_TARGET. The server reads configuration from the current working directory, not the repository root. ```bash hkex-scraper-mcp ``` -------------------------------- ### Adopting a new relational sink: configuration example Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Upgrading Example configuration for adding SQLite as a sink: sets DATABASE_TARGET to 'postgres,sqlite' and specifies SQLITE_PATH. ```ini DATABASE_TARGET=postgres,sqlite SQLITE_PATH=hkex.db ``` -------------------------------- ### Install per-sink extras Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Getting-Started Installs a single database driver extra. Replace postgres or duckdb with mysql, mongodb, clickhouse, or neo4j to add that driver. ```bash pip install "hkex-filing-scraper[postgres]" # add one driver at a time pip install "hkex-filing-scraper[duckdb]" # or: mysql, mongodb, clickhouse, neo4j ``` -------------------------------- ### Install a specific release from Git tag Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Releasing Install a specific release directly from the Git tag, useful for pinning to a particular version. ```bash pip install "git+https://github.com/simonplmak-cloud/hkex-filing-scraper@v1.2.0" ``` -------------------------------- ### Set up development environment and run checks Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/README.md Run these commands in the project root to set up the development environment and run checks. The pip install command installs the package in editable mode with development and all optional dependencies. The ruff commands lint and check formatting, while pytest runs unit tests that do not require a database or network. ```bash pip install -e ".[dev,all]" ruff check # lint (py310, line-length 100) ruff format --check # formatting pytest # unit tests (no DB or network required) ``` -------------------------------- ### Copy .env.example to .env Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Getting-Started Copies the example environment file to .env in the current working directory. The scraper loads Path.cwd()/.env, not the repo root. ```bash cp .env.example .env ``` -------------------------------- ### Start local database stack with Docker Compose Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/examples/README.md Starts the local database stack with Docker Compose. The comment lists the supported databases: postgres, mysql, mariadb, mongodb, clickhouse, neo4j, and surrealdb. Wait for health checks to settle using `docker compose ps`. ```bash cd examples docker compose up -d # postgres, mysql, mariadb, mongodb, clickhouse, neo4j, surrealdb ``` -------------------------------- ### Example SQL queries for exchange_filing and scrape_coverage Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/PostgreSQL Example SQL queries for the exchange_filing and scrape_coverage tables. Use them to analyze filings per company, find filings referencing a ticker, inspect extracted JSONB tables, check document status distribution, and review scrape coverage. ```sql -- Filings per company SELECT company_ticker, count(*) AS filings FROM exchange_filing GROUP BY company_ticker ORDER BY filings DESC LIMIT 20; -- Filings that reference a ticker SELECT filing_id, title, filing_date FROM exchange_filing WHERE referenced_tickers && ARRAY['0700.HK'] ORDER BY filing_date DESC; -- Inspect extracted tables (JSONB) SELECT filing_id, elem->>'pageNumber' AS page, left(elem->>'markdown', 200) AS preview FROM exchange_filing, jsonb_array_elements(document_tables) AS elem WHERE filing_type = 'Annual Report' LIMIT 5; -- Document processing status distribution SELECT document_status, count(*) FROM exchange_filing GROUP BY 1; -- Coverage report SELECT chunk_from, chunk_to, api_count, ingested_count, unique_count, run_id FROM scrape_coverage ORDER BY chunk_from DESC; ``` -------------------------------- ### Install MySQL extra Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/MySQL-and-MariaDB Install the mysql extra to enable the MySQL/MariaDB sink. Requires PyMySQL, a pure Python driver. ```bash pip install ".[mysql]" ``` -------------------------------- ### Install the package with pip Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/SQLite Installs the package from the current directory. No extra dependencies are required for the SQLite sink because it uses the Python standard library sqlite3 driver. ```bash pip install . ``` -------------------------------- ### Example ClickHouse queries Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/ClickHouse Example SQL queries for the ClickHouse sink, including counting filings, filtering by referenced tickers, and extracting document tables. ```sql SELECT company_ticker, count() FROM exchange_filing FINAL GROUP BY 1 ORDER BY 2 DESC LIMIT 10; SELECT filing_id, title FROM exchange_filing FINAL WHERE has(referenced_tickers, '0700.HK'); SELECT filing_id, JSONExtractString(table, 'markdown') FROM exchange_filing FINAL ARRAY JOIN JSONExtractArrayRaw(document_tables) AS table WHERE filing_type = 'Annual Report'; ``` -------------------------------- ### Migration example: DATABASE_TARGET before and after (SurrealDB + PostgreSQL) Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Upgrading Shows the old 'both' value replaced with 'surrealdb,postgres' for the 2.0.0 upgrade. ```ini # before DATABASE_TARGET=both # after (SurrealDB + PostgreSQL) DATABASE_TARGET=surrealdb,postgres ``` -------------------------------- ### Verify release with pip install and hkex-scraper --version Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Release-Automation Spot-check a release by installing the package directly from the git tag and verifying the CLI version. The wheel version must equal the tag. ```bash pip install "git+https://github.com/simonplmak-cloud/hkex-filing-scraper@v1.2.0" hkex-scraper --version ``` -------------------------------- ### Example SurrealQL queries for the SurrealDB sink Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/SurrealDB Example SurrealQL queries for the SurrealDB sink. Includes filings per company, filings referencing a ticker, and coverage data. ```surql -- Filings per company SELECT companyTicker, count() AS filings FROM exchange_filing GROUP BY companyTicker ORDER BY filings DESC LIMIT 20; -- Filings that reference a ticker SELECT filingId, title, filingDate FROM exchange_filing WHERE '0700.HK' INSIDE referencedTickers; -- Coverage SELECT chunkFrom, chunkTo, apiCount, ingestedCount, uniqueCount, runId FROM scrape_coverage ORDER BY chunkFrom DESC; ``` -------------------------------- ### Migration example: DATABASE_TARGET before (unset) and after (SurrealDB) Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Upgrading Shows the implicit default (unset) replaced with explicit 'surrealdb' for the 2.0.0 upgrade. ```ini # before (implicit default, SurrealDB) # DATABASE_TARGET was unset # after DATABASE_TARGET=surrealdb ``` -------------------------------- ### MongoDB URI with authentication Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/docs/sinks/mongodb.md Example MONGODB_URI with authentication credentials and authSource parameter. ```bash MONGODB_URI=mongodb://user:password@localhost:27017/?authSource=admin ``` -------------------------------- ### Run throwaway SurrealDB server with Docker Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/SurrealDB Starts a throwaway SurrealDB server for local testing. Uses the v3.2.4 image, binds to port 8000, and runs in memory with root credentials. ```bash docker run -d --name surrealdb -p 8000:8000 \ surrealdb/surrealdb:v3.2.4 start --log warn --user root --pass root --bind 0.0.0.0:8000 memory ``` -------------------------------- ### Configure SQLite sink with DATABASE_TARGET Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/README.md Configure the scraper to use SQLite as the only sink, with the database file at hkex.db. This setup requires no server and is the simplest way to start. ```ini DATABASE_TARGET=sqlite SQLITE_PATH=hkex.db ``` -------------------------------- ### Configure each sink in .env Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Getting-Started Example .env configuration for each supported sink. Set DATABASE_TARGET to the sink id and add that sink's connection settings. Order matters: reads come from the first read-capable sink in the list. ```ini # PostgreSQL DATABASE_TARGET=postgres POSTGRES_DSN=postgresql://user:password@localhost:5432/hkex # or: POSTGRES_HOST / POSTGRES_PORT / POSTGRES_DATABASE / POSTGRES_USER / POSTGRES_PASSWORD # MySQL / MariaDB DATABASE_TARGET=mysql # or mariadb MYSQL_HOST=localhost MYSQL_DATABASE=hkex MYSQL_USER=hkex MYSQL_PASSWORD=secret # SQLite (no server) DATABASE_TARGET=sqlite SQLITE_PATH=hkex.db # MongoDB DATABASE_TARGET=mongodb MONGODB_URI=mongodb://localhost:27017 MONGODB_DATABASE=hkex # Neo4j DATABASE_TARGET=neo4j NEO4J_URI=bolt://localhost:7687 NEO4J_USER=neo4j NEO4J_PASSWORD=secret # ClickHouse DATABASE_TARGET=clickhouse CLICKHOUSE_HOST=localhost CLICKHOUSE_DATABASE=hkex CLICKHOUSE_USER=default # DuckDB (no server) DATABASE_TARGET=duckdb DUCKDB_PATH=hkex.duckdb # SurrealDB DATABASE_TARGET=surrealdb SURREAL_ENDPOINT=http://localhost:8000 SURREAL_NAMESPACE=default SURREAL_DATABASE=default SURREAL_USERNAME=root SURREAL_PASSWORD=your_password ``` -------------------------------- ### Install ClickHouse extra Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/ClickHouse Install the ClickHouse extra dependency for the project. ```bash pip install ".[clickhouse]" ``` -------------------------------- ### Install DuckDB extra Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/DuckDB Installs the DuckDB sink extra. Required before using the DuckDB database target. ```bash pip install ".[duckdb]" ``` -------------------------------- ### Install PostgreSQL sink with pip Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/PostgreSQL Install the PostgreSQL sink extra, which includes psycopg[binary,pool]>=3.1. ```bash pip install ".[postgres]" # psycopg[binary,pool]>=3.1 ``` -------------------------------- ### Install MCP extra Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/MCP Install the package with the optional MCP extra using pip. The 'mcp' extra is not a base dependency, so it must be explicitly requested. ```bash pip install "hkex-filing-scraper[mcp]" ``` -------------------------------- ### Run multi-sink integration tests with environment variables Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Testing Sets environment variables for PostgreSQL, SurrealDB, and SQLite sinks, then runs the integration tests for those sinks. Every configured sink must have its schema initialized; the integration fixtures do this for the whole configured set, matching main._init_schemas(). ```bash export DATABASE_TARGET=postgres,surrealdb,sqlite export SURREAL_ENDPOINT=http://localhost:8000 export SURREAL_NAMESPACE=test export SURREAL_DATABASE=hkex export SURREAL_USERNAME=root export SURREAL_PASSWORD=root export POSTGRES_DSN=postgresql://hkex:hkex@localhost:5432/hkex export SQLITE_PATH=/tmp/hkex-it.db pytest -q tests/test_postgres_integration.py tests/test_surrealdb_integration.py tests/test_dual_write_integration.py ``` -------------------------------- ### Example Cypher queries Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Neo4j Example Cypher queries for common lookups: filings for a company, filings referencing another company, documents needing processing, and most-referenced companies. ```cypher // filings for one company MATCH (c:Company {id: '451_HK'})-[:HAS_FILING]->(f:Filing) RETURN f.title, f.filingDate ORDER BY f.filingDate DESC LIMIT 10; // filings that mention another company MATCH (f:Filing)-[:REFERENCES_FILING]->(c:Company {id: '700_HK'}) RETURN f.filingId, f.title; // documents that still need processing MATCH (f:Filing) WHERE f.documentStatus IS NULL AND f.documentUrl IS NOT NULL RETURN count(f); // most-referenced companies MATCH (:Filing)-[r:REFERENCES_FILING]->(c:Company) RETURN c.id, count(r) AS mentions ORDER BY mentions DESC LIMIT 10; ``` -------------------------------- ### Prompt for listing filings and summarising report Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/blob/main/README.md Example prompt for an AI agent using the hosted MCP gateway. It asks the agent to list filings in a date window and then summarise the interim report. ```text Use hkex-live to list the filings published between 2026-09-01 and 2026-09-18, then summarise the interim report. ``` -------------------------------- ### Run MongoDB, ClickHouse, and Neo4j integration tests Source: https://github.com/simonplmak-cloud/hkex-filing-scraper/wiki/Testing Runs the extra sinks integration suite for MongoDB, ClickHouse, and Neo4j. Requires the respective drivers and connection variables; otherwise the suite is skipped. ```bash export DATABASE_TARGET=mongodb,clickhouse,neo4j export MONGODB_URI=mongodb://localhost:27017 export MONGODB_DATABASE=hkex export CLICKHOUSE_HOST=localhost CLICKHOUSE_PORT=8123 CLICKHOUSE_DATABASE=hkex export CLICKHOUSE_USER=hkex CLICKHOUSE_PASSWORD=hkex export NEO4J_URI=bolt://localhost:7687 NEO4J_USER=neo4j NEO4J_PASSWORD=hkexpassword pip install -e ".[dev,duckdb,mongodb,clickhouse,neo4j]" pytest -q tests/test_extra_sinks_integration.py ```