|
Size: 1494
Comment:
|
Size: 4125
Comment:
|
| Deletions are marked like this. | Additions are marked like this. |
| Line 10: | Line 10: |
| 1. PitrBackup = 99% complete (currently working on a fix to the routine that updates/inserts the meta data) | 1. pitrBackup = 100% complete (We may still find bugs as we engage in further testing) |
| Line 12: | Line 12: |
| 2. walBackup = 75% complete, completion expected tomorrow (Thurs 9/10/2009) | 2. walBackup = 100% complete (We may still find bugs as we engage in further testing) |
| Line 16: | Line 16: |
| 4. Unit testing (first round for pitrBackup complete) | 4. Unit testing - 100% (Basic Unit tests) |
| Line 20: | Line 20: |
| Installed on hmidb | Installed on hmidb (after the move to hmidb0 we'll setup the final packages) |
| Line 22: | Line 22: |
| 6. Setup at least 2 cron jobs | 6. Setup backup jobs (pending move to hmidb0) |
| Line 26: | Line 26: |
| * not yet implemented, I'd like to wait for the hmidb0 setup to be in place | |
| Line 30: | Line 29: |
| Line 32: | Line 32: |
| * Documented strategy * Strategy fully dependant on the Backup / Recovery strategy and defined manual steps * We'll need to address the re-setup of SLONY after a fAIL over as part of the upcoming 'Data Replication' strategy |
* Documented strategy * Strategy fully dependant on the Backup / Recovery strategy and defined manual steps |
| Line 36: | Line 35: |
| * I successfully tested the backup of a running warm standby server and then recovering from the warm standby base backup and the subsequent archived WAL files from the master | * We'll need to address the re-setup of SLONY after a fAIL over as part of the upcoming 'Data Replication' strategy * I successfully tested the backup of a running warm standby server and then recovering from the warm standby base backup and the subsequent archived WAL files from the master * walBackup script 100% complete * walBackup Unit Testing - 100% complete (Basic unit tests) * Still To Do: Revise / make more clear & easy to follow the recovery plan |
| Line 39: | Line 46: |
| == Database Access and Security == | |
| Line 41: | Line 47: |
| * Starting this step tomorrow (09-10-2009) | == Web db Plan == * Strategy documented - initial pass (100%) == SLONY Plan == * SLONY v2 must be used if we want to use PostgreSQL 8.4 * SLONY-1 2.03 (release Candidate) recommended due to a few key bugs in 2.01 and 2.02 * Next Steps: * Define SLONY Architecture * Define Log Shipping process * Build SLONY scripts/tools |
| Line 43: | Line 65: |
| * Implement & Test == Currently In Progress == * Setup of a new (local) warm standby server (done) * further testing of the current backup scripts (done) * Design of the SLONY control modules (done) * Revise / make more clear & easy to follow the warm standby recovery plan (pending) * End2end testing: * setup 4 VM's (done) * install postgres on nodes 1 & 2 (done) * setup warm standby (done) * install PITR backup scripts (done) * install slony control scripts (done) * setup initial slony cluster (done, pending re-starts as needed) * test warm standby failover (done - first pass - success with caveats) * test slony switchover / switch back (done - success) * test pitr recovery (pending) * setup slony log shipping (pending) * test slony log shipping receiver (pending) * document end2end test results (pending) * add slony scripts to add/remove things from the slony cluster (pending) * test the add/remove things to slony script(s) (pending) * NOTES: we need to manage consistency of the pitr backups off the warm standby ourselves based on discussions with the Postgres development team I believe we should test the following: * shutdown the warm standby database(s), run the xfsdump, then restart (back into recovery mode) the warm standby database(s) * creates a risk factor in that the warm standby is down during the dump * Force a checkpoint on the master, watch the standby logs for the completion of the checkpoint and then run the xfsdump * creates the need to manage the 'window' between checkpoints and ensure that its long enough to do our xfsdump * This involves the monitoring/tweaking of the # of checkpoint segments and the checkpoint_timeout in relation to current traffic volues * Both of these, plus the single failover script has introduces unexpected issues and/or unexpected scope, thus pushing our timeline out. I'll try and make up some time in the next (monitoring) phase of the schedule. |
Kevins's archive
Implementation Status
Updated Wed 09-09-2009
Backup / Recovery
1. pitrBackup = 100% complete (We may still find bugs as we engage in further testing)
2. walBackup = 100% complete (We may still find bugs as we engage in further testing)
3. Implement changes based on feedback if needed
4. Unit testing - 100% (Basic Unit tests)
5. Install 'Package” as a directory structure that will contain all future 'tools' /bin, /etc, /tmp, /log, ... Installed on hmidb (after the move to hmidb0 we'll setup the final packages)
6. Setup backup jobs (pending move to hmidb0) Cron entry to run the base file system backup (I suggest once/quarter) Cron entry to archive the WAL segments (monthly) Optional additional base file system backups to alternate locations
The toolset will be run from the warm standby server
Warm Standby
- Documented strategy
- Strategy fully dependant on the Backup / Recovery strategy and defined manual steps
- We'll need to address the re-setup of SLONY after a fAIL over as part of the upcoming 'Data Replication' strategy
- I successfully tested the backup of a running warm standby server and then recovering from the warm standby base backup and the subsequent archived WAL files from the master
- walBackup script 100% complete
- walBackup Unit Testing - 100% complete (Basic unit tests)
Still To Do: Revise / make more clear & easy to follow the recovery plan
Web db Plan
- Strategy documented - initial pass (100%)
SLONY Plan
- SLONY v2 must be used if we want to use PostgreSQL 8.4
- SLONY-1 2.03 (release Candidate) recommended due to a few key bugs in 2.01 and 2.02
- Next Steps:
- Define SLONY Architecture
- Define Log Shipping process
- Build SLONY scripts/tools
Implement & Test
Currently In Progress
- Setup of a new (local) warm standby server (done)
- further testing of the current backup scripts (done)
- Design of the SLONY control modules (done)
Revise / make more clear & easy to follow the warm standby recovery plan (pending)
- End2end testing:
- setup 4 VM's (done)
install postgres on nodes 1 & 2 (done)
- setup warm standby (done)
- install PITR backup scripts (done)
- install slony control scripts (done)
- setup initial slony cluster (done, pending re-starts as needed)
- test warm standby failover (done - first pass - success with caveats)
- test slony switchover / switch back (done - success)
- test pitr recovery (pending)
- setup slony log shipping (pending)
- test slony log shipping receiver (pending)
- document end2end test results (pending)
- add slony scripts to add/remove things from the slony cluster (pending)
- test the add/remove things to slony script(s) (pending)
- NOTES: we need to manage consistency of the pitr backups off the warm standby ourselves based on discussions with the Postgres development team I believe we should test the following:
- shutdown the warm standby database(s), run the xfsdump, then restart (back into recovery mode) the warm standby database(s)
- creates a risk factor in that the warm standby is down during the dump
- Force a checkpoint on the master, watch the standby logs for the completion of the checkpoint and then run the xfsdump
- creates the need to manage the 'window' between checkpoints and ensure that its long enough to do our xfsdump
- This involves the monitoring/tweaking of the # of checkpoint segments and the checkpoint_timeout in relation to current traffic volues
- Both of these, plus the single failover script has introduces unexpected issues and/or unexpected scope, thus pushing our timeline out.
- I'll try and make up some time in the next (monitoring) phase of the schedule.
- shutdown the warm standby database(s), run the xfsdump, then restart (back into recovery mode) the warm standby database(s)
