Logfile Roundup
Logfile Roundup
@LogfileRoundup

Playbook: sample logs without lying to yourself

Playbook: sample logs without lying to yourself
When full logs won't fit in memory, these sources keep the math honest.
→ Google's own "sample your logs" guidance — for big sites, a representative slice beats a truncated one. Takeaway: sample by time window, never by truncating the file head.
→ Reservoir sampling explainer (Greg Reda) — pull a uniform random N lines in one pass. Takeaway: awk reservoir sampling lets you sample a 50GB file you can't even open.
★ Pick of the week — Hamlet Batista's stratified-sampling note — sample per status-code bucket so rare 5xx don't vanish. Takeaway: stratify by the dimension you actually care about, or your rare events disappear.
→ Pandas .sample(frac=) docs — quick once the data's already loaded. Takeaway: set a seed so results are reproducible.
State your sampling method in any report — unstated sampling is a silent bias.
Этот пост опубликован в Telegram-канале Logfile Roundup. Подписаться можно по ссылке: @LogfileRoundup.
tech

Свежие посты в категории «Tech Infrastructure»

Все каналы категории →

start

Готовы запустить рекламу через сеть public.tg?

Новый оффер, продукт, GEO, кейс, событие или партнёрский запуск — соберём маршрут под задачу и отдадим медиаплан.

Telegram для медиаплана: @AFFtop_connect. Быстрый тест: $20 за канал, $1000 за пакет по сети.