2022-02-28
ยง
|
10:02 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1126', diff saved to https://phabricator.wikimedia.org/P21578 and previous config saved to /var/cache/conftool/dbconfig/20220228-100221-ladsgroup.json |
[production] |
09:51 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Depooling db1179 (T300992)', diff saved to https://phabricator.wikimedia.org/P21577 and previous config saved to /var/cache/conftool/dbconfig/20220228-095056-ladsgroup.json |
[production] |
09:50 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance |
[production] |
09:50 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.downtime for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance |
[production] |
09:47 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1126', diff saved to https://phabricator.wikimedia.org/P21576 and previous config saved to /var/cache/conftool/dbconfig/20220228-094717-ladsgroup.json |
[production] |
09:32 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1126 (T302185)', diff saved to https://phabricator.wikimedia.org/P21575 and previous config saved to /var/cache/conftool/dbconfig/20220228-093212-ladsgroup.json |
[production] |
09:29 |
<volans@cumin1001> |
END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) |
[production] |
09:28 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1102.eqiad.wmnet with reason: Maintenance |
[production] |
09:28 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.downtime for 6:00:00 on db1102.eqiad.wmnet with reason: Maintenance |
[production] |
09:28 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1112 (T300992)', diff saved to https://phabricator.wikimedia.org/P21574 and previous config saved to /var/cache/conftool/dbconfig/20220228-092830-ladsgroup.json |
[production] |
09:27 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1126.eqiad.wmnet with OS bullseye |
[production] |
09:22 |
<volans@cumin1001> |
START - Cookbook sre.dns.netbox |
[production] |
09:16 |
<moritzm> |
restarting Hue to pick up expat security updates |
[production] |
09:13 |
<moritzm> |
restarting turnilo to pick up expat security updates |
[production] |
09:13 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1112', diff saved to https://phabricator.wikimedia.org/P21573 and previous config saved to /var/cache/conftool/dbconfig/20220228-091325-ladsgroup.json |
[production] |
09:12 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1126.eqiad.wmnet with reason: host reimage |
[production] |
09:10 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.downtime for 2:00:00 on db1126.eqiad.wmnet with reason: host reimage |
[production] |
09:00 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.reimage for host db1126.eqiad.wmnet with OS bullseye |
[production] |
08:58 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1112', diff saved to https://phabricator.wikimedia.org/P21572 and previous config saved to /var/cache/conftool/dbconfig/20220228-085820-ladsgroup.json |
[production] |
08:53 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Depooling db1126 (T302185)', diff saved to https://phabricator.wikimedia.org/P21571 and previous config saved to /var/cache/conftool/dbconfig/20220228-085329-ladsgroup.json |
[production] |
08:53 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1126.eqiad.wmnet with reason: Maintenance |
[production] |
08:53 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.downtime for 1 day, 0:00:00 on db1126.eqiad.wmnet with reason: Maintenance |
[production] |
08:52 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1177 (T302185)', diff saved to https://phabricator.wikimedia.org/P21570 and previous config saved to /var/cache/conftool/dbconfig/20220228-085224-ladsgroup.json |
[production] |
08:51 |
<moritzm> |
installing expat security updates |
[production] |
08:47 |
<ayounsi@cumin1001> |
END (FAIL) - Cookbook sre.dns.netbox (exit_code=99) |
[production] |
08:43 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1112 (T300992)', diff saved to https://phabricator.wikimedia.org/P21567 and previous config saved to /var/cache/conftool/dbconfig/20220228-084316-ladsgroup.json |
[production] |
08:39 |
<ayounsi@cumin1001> |
START - Cookbook sre.dns.netbox |
[production] |
08:37 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1177', diff saved to https://phabricator.wikimedia.org/P21566 and previous config saved to /var/cache/conftool/dbconfig/20220228-083720-ladsgroup.json |
[production] |
08:22 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1177', diff saved to https://phabricator.wikimedia.org/P21564 and previous config saved to /var/cache/conftool/dbconfig/20220228-082215-ladsgroup.json |
[production] |
08:10 |
<taavi> |
UTC morning deploys done |
[production] |
08:09 |
<taavi@deploy1002> |
Synchronized logos/config.yaml: Config: [[gerrit:766138|Change temporary logo for slwiki (T302661)]] (duration: 00m 48s) |
[production] |
08:09 |
<taavi@deploy1002> |
Synchronized wmf-config/logos.php: Config: [[gerrit:766138|Change temporary logo for slwiki (T302661)]] (duration: 00m 48s) |
[production] |
08:08 |
<taavi@deploy1002> |
Synchronized static/images/project-logos: Config: [[gerrit:766138|Change temporary logo for slwiki (T302661)]] (duration: 00m 50s) |
[production] |
08:07 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1177 (T302185)', diff saved to https://phabricator.wikimedia.org/P21563 and previous config saved to /var/cache/conftool/dbconfig/20220228-080710-ladsgroup.json |
[production] |
08:06 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Depooling db1112 (T300992)', diff saved to https://phabricator.wikimedia.org/P21562 and previous config saved to /var/cache/conftool/dbconfig/20220228-080613-ladsgroup.json |
[production] |
08:06 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb[1013,1017,1021].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance |
[production] |
08:06 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.downtime for 12:00:00 on clouddb[1013,1017,1021].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance |
[production] |
08:06 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1112.eqiad.wmnet with reason: Maintenance |
[production] |
08:06 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.downtime for 6:00:00 on db1112.eqiad.wmnet with reason: Maintenance |
[production] |
08:06 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1123 (T300992)', diff saved to https://phabricator.wikimedia.org/P21561 and previous config saved to /var/cache/conftool/dbconfig/20220228-080559-ladsgroup.json |
[production] |
08:01 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1177.eqiad.wmnet with OS bullseye |
[production] |
08:00 |
<godog> |
enable notifications for thanos-be1003 in icinga and clear up /srv/swift-storage/sdm1 since it was filling up / |
[production] |
07:58 |
<moritzm> |
drain instances off ganeti2007 for eventual decom |
[production] |
07:50 |
<ladsgroup@cumin1001> |
dbctl commit (dc=all): 'Repooling after maintenance db1123', diff saved to https://phabricator.wikimedia.org/P21560 and previous config saved to /var/cache/conftool/dbconfig/20220228-075054-ladsgroup.json |
[production] |
07:45 |
<ladsgroup@cumin1001> |
END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1177.eqiad.wmnet with reason: host reimage |
[production] |
07:45 |
<elukey@deploy1002> |
helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'. |
[production] |
07:44 |
<elukey@deploy1002> |
helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'. |
[production] |
07:43 |
<elukey@deploy1002> |
helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'. |
[production] |
07:42 |
<ladsgroup@cumin1001> |
START - Cookbook sre.hosts.downtime for 2:00:00 on db1177.eqiad.wmnet with reason: host reimage |
[production] |
07:42 |
<elukey@deploy1002> |
helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'. |
[production] |