Logfile Roundup
Logfile Roundup
@LogfileRoundup

Playbook: stand up a BigQuery log pipeline

Playbook: stand up a BigQuery log pipeline
Four references for moving from grep-on-a-box to queryable logs.
→ Google Cloud's log sink docs — route Cloud Logging straight to a BigQuery dataset with a single sink. Takeaway: partition by ingestion date or your scans get expensive fast.
→ Distilled / Builtvisible's classic log-to-BigQuery write-up — the schema for parsing Apache lines into typed columns. Takeaway: store status and bytes as INT64, not strings.
★ Pick of the week — JR Oakes' SQL for crawl frequency by URL — a GROUP BY url, DATE that gives days-between-crawls per page. Takeaway: this single query replaces a week of spreadsheet work.
→ dbt log-parsing example repos — regex extraction in a model so transforms stay version-controlled. Takeaway: keep the UA-parsing logic in one place.
Partition and cluster up front — retrofitting a 100GB table hurts.
Этот пост опубликован в Telegram-канале Logfile Roundup. Подписаться можно по ссылке: @LogfileRoundup.
tech

Свежие посты в категории «Tech Infrastructure»

Все каналы категории →

start

Готовы запустить рекламу через сеть public.tg?

Новый оффер, продукт, GEO, кейс, событие или партнёрский запуск — соберём маршрут под задачу и отдадим медиаплан.

Telegram для медиаплана: @AFFtop_connect. Быстрый тест: $20 за канал, $1000 за пакет по сети.