1-50 of 10000 results (27ms)
2026-08-18 ยง
23:38 <ladsgroup@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326944|Retire filebackend lock manager in favour of the default one (T366938)]] (duration: 08m 55s) [production]
23:34 <ladsgroup@deploy1003> ladsgroup: Continuing with deployment [production]
23:31 <ladsgroup@deploy1003> ladsgroup: Backport for [[gerrit:1326944|Retire filebackend lock manager in favour of the default one (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
23:29 <ladsgroup@deploy1003> Started scap sync-world: Backport for [[gerrit:1326944|Retire filebackend lock manager in favour of the default one (T366938)]] [production]
23:27 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1170.eqiad.wmnet with OS bookworm [production]
23:21 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1205.eqiad.wmnet with OS bookworm [production]
23:20 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1171.eqiad.wmnet with OS bookworm [production]
23:15 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-worker1206.eqiad.wmnet with OS bookworm [production]
23:05 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage [production]
23:00 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage [production]
22:57 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage [production]
22:54 <ryankemper@cumin2003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage [production]
22:53 <ryankemper@cumin2003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1205.eqiad.wmnet with reason: host reimage [production]
22:51 <ryankemper@cumin2003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1171.eqiad.wmnet with reason: host reimage [production]
22:51 <ryankemper@cumin2003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1170.eqiad.wmnet with reason: host reimage [production]
22:50 <ryankemper@cumin2003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-worker1206.eqiad.wmnet with reason: host reimage [production]
22:36 <ryankemper@cumin2003> START - Cookbook sre.hosts.reimage for host an-worker1206.eqiad.wmnet with OS bookworm [production]
22:35 <ryankemper@cumin2003> START - Cookbook sre.hosts.reimage for host an-worker1205.eqiad.wmnet with OS bookworm [production]
22:35 <ryankemper@cumin2003> START - Cookbook sre.hosts.reimage for host an-worker1186.eqiad.wmnet with OS bookworm [production]
22:35 <ryankemper@cumin2003> START - Cookbook sre.hosts.reimage for host an-worker1171.eqiad.wmnet with OS bookworm [production]
22:35 <ryankemper@cumin2003> START - Cookbook sre.hosts.reimage for host an-worker1170.eqiad.wmnet with OS bookworm [production]
22:33 <ryankemper@cumin2003> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-worker1194.eqiad.wmnet with OS bookworm [production]
22:22 <ladsgroup@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326923|Enable redis lock manager on s6 (T366938)]] (duration: 11m 52s) [production]
22:18 <ladsgroup@deploy1003> ladsgroup: Continuing with deployment [production]
22:12 <ladsgroup@deploy1003> ladsgroup: Backport for [[gerrit:1326923|Enable redis lock manager on s6 (T366938)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
22:10 <ladsgroup@deploy1003> Started scap sync-world: Backport for [[gerrit:1326923|Enable redis lock manager on s6 (T366938)]] [production]
22:04 <sbassett> Deployed security fix for T435234 (wmf.16) [production]
21:54 <sbassett> Deployed security fix for T435234 (wmf.15) [production]
21:38 <caro@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326925|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] (duration: 0 [production]
21:34 <caro@deploy1003> caro: Continuing with deployment [production]
21:33 <caro@deploy1003> caro: Backport for [[gerrit:1326925|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] synced to the testservers (see h [production]
21:31 <caro@deploy1003> Started scap sync-world: Backport for [[gerrit:1326925|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326926|Exclude NPOV-type LLM-generated suggestions (T435253)]], [[gerrit:1326929|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]], [[gerrit:1326928|LLMSuggestionsEditCheck: final comparison should also have the object-replacements]] [production]
21:24 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: 1204 datanode repair T434494 [production]
21:07 <eileen> civicrm upgraded from 49047af1 to 08ca9d6c [fundraising]
21:02 <krinkle@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] (duration: 13m 56s) [production]
20:58 <krinkle@deploy1003> krinkle: Continuing with deployment [production]
20:50 <krinkle@deploy1003> krinkle: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
20:49 <ryankemper> `an-launcher1003` terminated process group `666809` (`rest_backfill_phase1.sh`) ~20 mins ago with `sudo kill -TERM -- -666809` after its local spark driver (`--driver-memory 64g`) repeatedly exhausted memory on the 32 GB VM and caused SSH to intermittently flap; host recovered to 27 GB available memory [production]
20:48 <krinkle@deploy1003> Started scap sync-world: Backport for [[gerrit:1326896|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]], [[gerrit:1326897|ext.math.mathjax: Call MathJax.typeset() from `wikipage.content` hook (T434469 T419356 T422077)]] [production]
20:35 <ladsgroup@deploy1003> helmfile [codfw] DONE helmfile.d/services/thumbor: apply [production]
20:33 <ladsgroup@deploy1003> helmfile [codfw] START helmfile.d/services/thumbor: apply [production]
20:31 <ladsgroup@deploy1003> helmfile [eqiad] DONE helmfile.d/services/thumbor: apply [production]
20:31 <kemayo@deploy1003> Finished scap sync-world: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] (duration: 07m 35s) [production]
20:28 <ladsgroup@deploy1003> helmfile [eqiad] START helmfile.d/services/thumbor: apply [production]
20:26 <kemayo@deploy1003> kemayo: Continuing with deployment [production]
20:26 <ryankemper> `an-launcher1003` confirmed the host is flapping because of memory thrash. chasing down the source of the thrash [production]
20:25 <kemayo@deploy1003> kemayo: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
20:23 <kemayo@deploy1003> Started scap sync-world: Backport for [[gerrit:1326870|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]], [[gerrit:1326871|LLMSuggestionsEditCheck: don't over-cache the description (T428641)]] [production]
20:23 <ladsgroup@deploy1003> helmfile [staging] DONE helmfile.d/services/thumbor: apply [production]
20:20 <ladsgroup@deploy1003> helmfile [staging] START helmfile.d/services/thumbor: apply [production]