|
2026-07-20
ยง
|
| 15:58 |
<bking@cumin2003> |
END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430880, restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards |
[production] |
| 15:53 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Depooling db2242 (T431660)', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json |
[production] |
| 15:53 |
<cwilliams@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance |
[production] |
| 15:44 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Repooling after maintenance db2162 (T431660)', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json |
[production] |
| 15:35 |
<sukhe@cumin1003> |
cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet |
[production] |
| 15:34 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json |
[production] |
| 15:33 |
<sukhe@cumin1003> |
cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet |
[production] |
| 15:28 |
<bking@cumin2003> |
START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - T431826 |
[production] |
| 15:24 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json |
[production] |
| 15:14 |
<bking@cumin2003> |
END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 |
[production] |
| 15:14 |
<bking@cumin2003> |
START - Cookbook sre.hosts.move-vlan for host wdqs1023 |
[production] |
| 15:14 |
<bking@cumin2003> |
START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm |
[production] |
| 15:14 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Repooling after maintenance db2162 (T431660)', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json |
[production] |
| 15:13 |
<urbanecm@deploy2003> |
Finished scap sync-world: Backport for [[gerrit:1312471|postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423|Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) |
[production] |
| 15:08 |
<kevinbazira@deploy2003> |
helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . |
[production] |
| 15:07 |
<cwilliams@cumin1003> |
dbctl commit (dc=all): 'Depooling db2162 (T431660)', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json |
[production] |
| 15:07 |
<cwilliams@cumin1003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance |
[production] |
| 15:05 |
<bking@cumin2003> |
END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430880, restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards |
[production] |
| 15:00 |
<urbanecm@deploy2003> |
vadymts1, migr, urbanecm: Continuing with deployment |
[production] |
| 14:59 |
<bking@cumin2003> |
END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - T431826 |
[production] |
| 14:58 |
<bking@deploy2003> |
Finished deploy [wdqs/wdqs@e8fb00c]: T430880 (duration: 00m 07s) |
[production] |
| 14:58 |
<bking@deploy2003> |
Started deploy [wdqs/wdqs@e8fb00c]: T430880 |
[production] |
| 14:58 |
<bking@deploy2003> |
Finished deploy [wdqs/wdqs@e8fb00c]: T430880 (duration: 00m 13s) |
[production] |
| 14:58 |
<bking@deploy2003> |
Started deploy [wdqs/wdqs@e8fb00c]: T430880 |
[production] |
| 14:57 |
<bking@cumin2003> |
START - Cookbook sre.wdqs.data-transfer (T430880, restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards |
[production] |
| 14:57 |
<sukhe@cumin1003> |
cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet |
[production] |
| 14:55 |
<sukhe@cumin1003> |
cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet |
[production] |
| 14:51 |
<bking@cumin2003> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm |
[production] |
| 14:49 |
<sukhe@cumin1003> |
END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P{lvs7003*} and A:liberica |
[production] |
| 14:47 |
<urbanecm@deploy2003> |
vadymts1, migr, urbanecm: Backport for [[gerrit:1312471|postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423|Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. |
[production] |
| 14:44 |
<sukhe@cumin1003> |
END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] |
[production] |
| 14:44 |
<sukhe@cumin1003> |
START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] |
[production] |
| 14:41 |
<sukhe@cumin1003> |
END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P{lvs7003.magru.wmnet} and A:liberica |
[production] |
| 14:41 |
<sukhe@cumin1003> |
END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs7003.magru.wmnet} and A:liberica |
[production] |
| 14:41 |
<sukhe@cumin1003> |
START - Cookbook sre.loadbalancer.admin pooling P{lvs7003.magru.wmnet} and A:liberica |
[production] |
| 14:41 |
<sukhe@cumin1003> |
END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs7003.magru.wmnet} and A:liberica |
[production] |
| 14:39 |
<sukhe@cumin1003> |
START - Cookbook sre.loadbalancer.admin depooling P{lvs7003.magru.wmnet} and A:liberica |
[production] |
| 14:39 |
<sukhe@cumin1003> |
START - Cookbook sre.loadbalancer.upgrade restart P{lvs7003.magru.wmnet} and A:liberica |
[production] |
| 14:33 |
<sukhe@cumin1003> |
END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] |
[production] |
| 14:33 |
<sukhe@cumin1003> |
START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] |
[production] |
| 14:31 |
<urbanecm@deploy2003> |
Started scap sync-world: Backport for [[gerrit:1312471|postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423|Modify user groups rights in English Wikiquote (T432557)]] |
[production] |
| 14:24 |
<jgiannelos@deploy2003> |
helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply |
[production] |
| 14:24 |
<jgiannelos@deploy2003> |
helmfile [codfw] START helmfile.d/services/mw-parsoid: apply |
[production] |
| 14:24 |
<jgiannelos@deploy2003> |
helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply |
[production] |
| 14:24 |
<jgiannelos@deploy2003> |
helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply |
[production] |
| 14:23 |
<bking@cumin2003> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage |
[production] |
| 14:22 |
<bking@cumin2003> |
START - Cookbook sre.wdqs.data-transfer (T430880, restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards |
[production] |
| 14:19 |
<bking@cumin2003> |
START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage |
[production] |
| 14:19 |
<bking@cumin2003> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm |
[production] |
| 14:16 |
<sukhe@cumin1003> |
cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet |
[production] |