51-100 of 10000 results (172ms)
2026-08-18 ยง
21:24 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair T434494 [production]
21:02 <krinkle@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) [production]
20:58 <krinkle@deploy1003> krinkle: Continuing with deployment [production]
20:50 <krinkle@deploy1003> krinkle: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
20:49 <ryankemper> `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory [production]
20:48 <krinkle@deploy1003> Started scap sync-world: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] [production]
20:35 <ladsgroup@deploy1003> helmfile [codfw] DONE helmfile.d/services/thumbor: apply [production]
20:33 <ladsgroup@deploy1003> helmfile [codfw] START helmfile.d/services/thumbor: apply [production]
20:31 <ladsgroup@deploy1003> helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [production]
20:31 <kemayo@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) [production]
20:28 <ladsgroup@deploy1003> helmfile [eqiad] START helmfile.d/services/thumbor: apply [production]
20:26 <kemayo@deploy1003> kemayo: Continuing with deployment [production]
20:26 <ryankemper> `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash [production]
20:25 <kemayo@deploy1003> kemayo: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
20:23 <kemayo@deploy1003> Started scap sync-world: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] [production]
20:23 <ladsgroup@deploy1003> helmfile [staging] DONE helmfile.d/services/thumbor: apply [production]
20:20 <ladsgroup@deploy1003> helmfile [staging] START helmfile.d/services/thumbor: apply [production]
20:14 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1168.eqiad.wmnet with OS bookworm [production]
20:08 <zabe> zabe@deploy1003:~$ mwscript extensions/WikimediaMaintenance/maintenance/fixFileRevisionArchiveNameDrift.php enwiki # T428406 [production]
20:08 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1204.eqiad.wmnet with OS bookworm [production]
20:05 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1167.eqiad.wmnet with OS bookworm [production]
20:02 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1166.eqiad.wmnet with OS bookworm [production]
19:54 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1203.eqiad.wmnet with OS bookworm [production]
19:53 <ladsgroup@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326907|Revert "Disable redis lock manager on testwiki"]] (duration: 11m 05s) [production]
19:50 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage [production]
19:47 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage [production]
19:46 <ladsgroup@deploy1003> ladsgroup: Continuing with deployment [production]
19:44 <ladsgroup@deploy1003> ladsgroup: Backport for [[gerrit:1326907|Revert "Disable redis lock manager on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
19:42 <ladsgroup@deploy1003> Started scap sync-world: Backport for [[gerrit:1326907|Revert "Disable redis lock manager on testwiki"]] [production]
19:41 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage [production]
19:37 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage [production]
19:34 <btullis@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage [production]
19:32 <btullis@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1167.eqiad.wmnet with reason: host reimage [production]
19:32 <btullis@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1168.eqiad.wmnet with reason: host reimage [production]
19:32 <btullis@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1166.eqiad.wmnet with reason: host reimage [production]
19:31 <btullis@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1204.eqiad.wmnet with reason: host reimage [production]
19:31 <btullis@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1203.eqiad.wmnet with reason: host reimage [production]
19:26 <reedy@deploy1003> Finished scap sync-world: Backport for [[gerrit:1324751|InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875|Add banner notifying of upcoming 2FA enforcement (T420792)]] (duration: 31m 46s) [production]
19:16 <btullis@cumin1003> START - Cookbook sre.hosts.reimage for host an-worker1204.eqiad.wmnet with OS bookworm [production]
19:16 <btullis@cumin1003> START - Cookbook sre.hosts.reimage for host an-worker1203.eqiad.wmnet with OS bookworm [production]
19:16 <btullis@cumin1003> START - Cookbook sre.hosts.reimage for host an-worker1168.eqiad.wmnet with OS bookworm [production]
19:16 <btullis@cumin1003> START - Cookbook sre.hosts.reimage for host an-worker1167.eqiad.wmnet with OS bookworm [production]
19:16 <btullis@cumin1003> START - Cookbook sre.hosts.reimage for host an-worker1166.eqiad.wmnet with OS bookworm [production]
19:15 <denisse> rebooting kafkamon2003.codfw.wmnet - T435162 [production]
19:14 <denisse> rebooting kafkamon1003.eqiad.wmnet T435162 [production]
19:13 <reedy@deploy1003> reedy: Continuing with deployment [production]
19:12 <reedy@deploy1003> reedy: Backport for [[gerrit:1324751|InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875|Add banner notifying of upcoming 2FA enforcement (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
18:54 <reedy@deploy1003> Started scap sync-world: Backport for [[gerrit:1324751|InitialiseSettings: Enable 2FA enforcement on more private wikis (T428103)]], [[gerrit:1326875|Add banner notifying of upcoming 2FA enforcement (T420792)]] [production]
18:50 <jasmine@cumin2002> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl2006.codfw.wmnet with OS trixie [production]
18:44 <aklapper@deploy1003> rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.16 refs T430835 [production]