351-400 of 10000 results (150ms)
2026-07-29 ยง
13:59 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance [production]
13:59 <root@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance [production]
13:58 <sukhe@puppetserver1001> conftool action : set/pooled=true; selector: dnsdisc=urldownloader [production]
13:58 <root@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie [production]
13:56 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db1218 (T431660)', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json [production]
13:56 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance [production]
13:55 <root@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance [production]
13:54 <mvernon@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie [production]
13:53 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db1233 (T431660)', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json [production]
13:53 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance [production]
13:53 <root@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance [production]
13:50 <jgiannelos@deploy1003> helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply [production]
13:50 <sukhe> sukhe@lvs2013:~$ sudo systemctl restart pybal.service [production]
13:49 <blake@deploy1003> helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [production]
13:49 <jgiannelos@deploy1003> helmfile [eqiad] START helmfile.d/services/kartotherian: apply [production]
13:48 <jgiannelos@deploy1003> helmfile [codfw] DONE helmfile.d/services/kartotherian: apply [production]
13:47 <mvernon@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie [production]
13:47 <jgiannelos@deploy1003> helmfile [codfw] START helmfile.d/services/kartotherian: apply [production]
13:46 <swfrench@cumin2002> START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm [production]
13:44 <jgiannelos@deploy1003> helmfile [codfw] DONE helmfile.d/services/kartotherian: apply [production]
13:44 <root@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage [production]
13:40 <stran@deploy1003> Finished scap sync-world: Backport for [[gerrit:1319079|SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080|SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077|SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078|SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s) [production]
13:39 <blake@deploy1003> helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [production]
13:38 <blake@deploy1003> helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [production]
13:36 <root@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage [production]
13:35 <stran@deploy1003> stran: Continuing with deployment [production]
13:33 <jgiannelos@deploy1003> helmfile [codfw] START helmfile.d/services/kartotherian: apply [production]
13:32 <stran@deploy1003> stran: Backport for [[gerrit:1319079|SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080|SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077|SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078|SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t [production]
13:32 <mvernon@cumin2003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage [production]
13:30 <klausman@cumin1003> END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet [production]
13:30 <stran@deploy1003> Started scap sync-world: Backport for [[gerrit:1319079|SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080|SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077|SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078|SI: Instrument abuse filter hits link (T433053)]] [production]
13:29 <sukhe@cumin1003> END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad [production]
13:28 <mvernon@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage [production]
13:27 <sukhe@cumin1003> START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad [production]
13:27 <blake@deploy1003> helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [production]
13:27 <sukhe@cumin1003> END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad [production]
13:26 <mvernon@cumin2003> START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage [production]
13:26 <samtar@deploy1003> Finished scap sync-world: Backport for [[gerrit:1309675|Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s) [production]
13:25 <klausman@cumin1003> START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet [production]
13:24 <sukhe@cumin1003> START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad [production]
13:24 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P{ml-serve101[2-5].eqiad.wmnet} and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad) [production]
13:24 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet [production]
13:24 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet [production]
13:24 <mvernon@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage [production]
13:23 <root@cumin1003> END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265 [production]
13:23 <root@cumin1003> START - Cookbook sre.hosts.move-vlan for host db1265 [production]
13:23 <root@cumin1003> START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie [production]
13:23 <root@cumin1003> START - Cookbook sre.mysql.pool pool db1243: Maintenance [production]
13:22 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply [production]
13:22 <swfrench@cumin2002> END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet [production]