/ - Changes - BIEN 3 - NCEAS Projects

root @ 10038

#	Date	Author	Comment
10038	06/25/2013 03:46 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: upload(): .rsync_filter: also support machine-specific filters, for cases when different machines produce the same file (e.g. a log file) but only one machine's copy should be backed up
10037	06/25/2013 03:43 PM	Aaron Marcuse-Kubitza	/README.TXT: Maintenance: to synchronize a Mac's settings with my testing machine's: removed filters that are now handled by .rsync_ignores
10036	06/25/2013 03:31 PM	Aaron Marcuse-Kubitza	added inputs/GBIF/_src/.rsync_filter.upload,download to prevent old versions of GBIFPortalDB-*.dump.gz from being downloaded to the local machine, while keeping them on jupiter. this avoids the need to store these files in ~/Documents/BIEN/large_files/ with symlinks from inputs/GBIF/_src/ to exclude them from the sync.
10035	06/25/2013 03:17 PM	Aaron Marcuse-Kubitza	bugfix: /README.TXT: Maintenance: to synchronize a Mac's settings with my testing machine's: sync ~/Dropbox/svn/ (the no-unversioned-files working copy) separately from the rest of the files, because .svn/ is now excluded by /.rsync_ignore, so that `svn up` needs to be used to keep the .svn/ dirs in sync. note that .svn/ should generally not be synced between machines, because they may use incompatible versions of the svn working copy format.
10034	06/25/2013 03:02 PM	Aaron Marcuse-Kubitza	/README.TXT: Maintenance: to synchronize a Mac's settings with my testing machine's: use new bin/sync_upload (with $sync_remote_subdir) so that per-dir .rsync_ignores are processed, and to use the default $sync_remote_url
10033	06/25/2013 02:57 PM	Aaron Marcuse-Kubitza	lib/sh/local.sh: $sync_remote_url: allow user to override just the sync subdir (not the whole URL) in $sync_remote_subdir. this is useful e.g. for backing up the Mac's files to jupiter.
10032	06/25/2013 02:28 PM	Aaron Marcuse-Kubitza	/README.TXT: Maintenance: to synchronize vegbiendev, jupiter, and your local machine: use new bin/sync_upload instead of specifying all the filter patterns manually. this replaces several `put` commands with various filters with just a bin/sync_upload each on vegbiendev and your machine (in overwrite=1 mode to force a complete sync).
10031	06/25/2013 02:21 PM	Aaron Marcuse-Kubitza	bugfix: backups/.rsync_filter.download: need to prevent existing backups from being deleted on the local side, too, by changing hide patterns to exclude
10030	06/25/2013 02:11 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: upload(): make put's $subpath option relative to the currdir instead, like the --include paths. note that $subpath unfortunately can't be used in subdirs at this point because it will cause rsync to ignore the .rsync_ignores and .rsync_filters in parent dirs, including the essential .rsync_ignore in the sync root dir.
10029	06/25/2013 01:42 PM	Aaron Marcuse-Kubitza	/README.TXT: removed unnecessary `env` before kw params, which are treated as such whenever they appear before a command name
10028	06/25/2013 01:22 PM	Aaron Marcuse-Kubitza	bugfix: /README.TXT: updated `make backups/download` to `make backups/<file>/download`
10027	06/25/2013 01:21 PM	Aaron Marcuse-Kubitza	backups/Makefile: upload: use bin/sync_upload
10026	06/25/2013 01:12 PM	Aaron Marcuse-Kubitza	inputs/Makefile: download-logs: use bin/sync_upload like upload/download
10025	06/25/2013 01:07 PM	Aaron Marcuse-Kubitza	bugfix: /README.TXT: `make inputs/upload`, `make inputs/download`: added live=1 so that the sync operation runs rather than previewing what will be synced. removed test=1 because this flag is not used by put.
10024	06/25/2013 01:00 PM	Aaron Marcuse-Kubitza	bugfix: inputs/Makefile: upload, download: need to exclude files in .rsync_ignore, so that large local-only files, such as inputs/GBIF/raw_occurrence_record_plants/table*.tsv, do not have to be synced before `make inputs/upload` can complete (the corresponding .gz gets extracted instead); and deleted temp files in inputs/VegBIEN/TWiki/, such as active sessions, are not added back to the live copy on vegbiendev. previously, fixing this required extracting the rsync command run by `make inputs/upload`, etc. and manually editing it to exclude the applicable .rsync_ignore files, each time `make inputs/upload`, etc. was run (including before every column-based import).
10023	06/25/2013 12:23 PM	Aaron Marcuse-Kubitza	bugfix: bin/make: need to leave bin/, ~/bin/ in the PATH when running make nonrecursively, so that commands invoked by it which are located in these dirs (e.g. put, which will be used by `make inputs/upload`) can still be found. this requires using command()'s new nonrecursive=1 flag instead of running no_PATH_recursion, so that no_PATH_recursion() only affects the resolution of the command path, but does not propagate the filtered PATH to the invoked command itself.
10022	06/25/2013 12:18 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: command(): added nonrecursive=1 flag, which uses cmd2abs_path to run an external command nonrecursively
10021	06/25/2013 12:16 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: added cmd2abs_path, which makes the command in $1 nonrecursive
10020	06/25/2013 11:37 AM	Aaron Marcuse-Kubitza	bugfix: lib/sh/util.sh: PATH_rm(): also need to remove adjacent occurrences of the same path (or occurrences which become adjacent when other paths are removed), which :...: matching wasn't doing because the trailing : is consumed, preventing it from being matched at the beginning of the next path. since unlike filesystem paths with /, it is not necessary for a match to span multiple :-separated sections, we can just use new split() to split apart the PATH into an array of paths, filter each path, and join() them back together.
10019	06/25/2013 11:33 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: added split()
10018	06/25/2013 10:32 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: auto-echo common external commands: added `which`
10017	06/25/2013 10:32 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: auto-echo common external commands: use simpler echo_run instead of command since logging handling is not needed
10016	06/25/2013 08:50 AM	Aaron Marcuse-Kubitza	added backups/vegbien.r9897.backup.md5
10015	06/23/2013 09:30 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: upload(): documented that each --include path is relative to the currdir, not the root dir of the upload ($local_dir). this feature, although previously unintended, is actually better because the user can change to a subdir of the root dir and specify upload paths relative to the dir they are in. however, when invoking upload() from a script with --include paths specified, this means you need to use an absolute path (e.g. "$(dirname "${BASH_SOURCE⁰}")"/...; or the value that will become $local_dir, which for sync_upload() is $root_dir).
10014	06/23/2013 08:56 PM	Aaron Marcuse-Kubitza	backups/.rsync_ignore: replaced with .rsync_filter.upload to allow uploading new backups but not deleting existing backups if they don't exist on the local (rsync-invoking) side; and .rsync_filter.download to avoid downloading backups to the local side. this allows storing older backups just on jupiter, where there is much more disk space. note that this change must be made on the remote side (jupiter) for it to be effective, because these are remote-side rules and are only processed by the remote-side rsync instance.
10013	06/23/2013 08:55 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: upload(): use directional .rsync_filter to supplement .rsync_ignore with all kinds of --filter rules. separate .rsync_filters are needed for the upload (swap=) and download (swap=1) directions because the sender and the receiver are reversed, causing asymmetric rules like protect/hide to change meaning.
10012	06/23/2013 07:48 PM	Aaron Marcuse-Kubitza	updated backups/TNRS.backup.md5
10011	06/23/2013 07:48 PM	Aaron Marcuse-Kubitza	added backups/TNRS.2013-6-17.backup.md5, TNRS.2013-6-22.backup.md5
10010	06/23/2013 03:58 PM	Aaron Marcuse-Kubitza	/README.TXT: Backups: TNRS cache: Back up/Restore: added runtimes (3 min/5.5 min)
10009	06/23/2013 03:52 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: upload(): usage: documented put's swap=1 flag, which downloads instead of uploads
10008	06/23/2013 03:47 PM	Aaron Marcuse-Kubitza	added inputs/GBIF/raw_occurrence_record_plants/.rsync_ignore with filters that have previously needed to be manually added whenever `make inputs/upload` was run
10007	06/23/2013 03:46 PM	Aaron Marcuse-Kubitza	added inputs/GBIF/_MySQL/.rsync_ignore with filters from /README.TXT > Maintenance > to synchronize vegbiendev, jupiter, and your local machine. these filters will now be used with bin/sync_upload in addition to the periodic backup commands.
10006	06/23/2013 03:45 PM	Aaron Marcuse-Kubitza	added inputs/VegBIEN/TWiki/.rsync_ignore with filters from /README.TXT > Maintenance > to synchronize vegbiendev, jupiter, and your local machine. these filters will now be used with bin/sync_upload in addition to the periodic backup commands.
10005	06/23/2013 03:44 PM	Aaron Marcuse-Kubitza	added inputs/.rsync_ignore with filters from inputs/Makefile $(rsyncSrcs). these filters will now be used with bin/sync_upload in addition to `make inputs/upload`.
10004	06/23/2013 03:43 PM	Aaron Marcuse-Kubitza	added bin/.rsync_ignore with filters from /README.TXT > Maintenance > to synchronize vegbiendev, jupiter, and your local machine. these filters will now be used with bin/sync_upload in addition to the periodic backup commands.
10003	06/23/2013 03:40 PM	Aaron Marcuse-Kubitza	added backups/.rsync_ignore with filters from /README.TXT > Maintenance > to synchronize vegbiendev, jupiter, and your local machine. these filters will now be used with bin/sync_upload in addition to the periodic backup commands.
10002	06/23/2013 03:36 PM	Aaron Marcuse-Kubitza	/.rsync_ignore: added *.pyc
10001	06/23/2013 03:36 PM	Aaron Marcuse-Kubitza	added /.rsync_ignore with filters from lib/common.Makefile $(rsync). these filters will now be used with bin/sync_upload in addition to `make inputs/upload`.
10000	06/23/2013 03:34 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: upload(): use --exclude filters from per-dir .rsync_ignore. note that --exclude-from can't be used for this, because it is relative to the currdir, not the rsync root, and therefore also requires the .rsync_ignore to exist rather than using it only if it exists.
9999	06/22/2013 12:23 AM	Aaron Marcuse-Kubitza	bin/tnrs_db: documented total runtime (10 days)
9998	06/21/2013 11:58 PM	Aaron Marcuse-Kubitza	bin/tnrs_db: documented current runtime (162 ms/name)
9997	06/20/2013 10:58 PM	Aaron Marcuse-Kubitza	web/links/index.htm: updated to Firefox bookmarks. sorted NCEAS bookmarks to put homepage and support pages first.
9996	06/20/2013 06:21 PM	Aaron Marcuse-Kubitza	/README.TXT: Full database import: To run TNRS, etc. after the main import: clarified that you should only run `export version=<version>` if the import is named something other than public (i.e. it has not yet replaced the previous public schema)
9995	06/20/2013 06:14 PM	Aaron Marcuse-Kubitza	/README.TXT: Full database import: To run TNRS: removed `by_col=1` because by-column mode is not applicable to running TNRS. it is, however, needed when running import_scrub (i.e. `make inputs/<datasrc>/reimport_scrub by_col=1`).
9994	06/20/2013 06:10 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs: vegbiendev update steps: added `make backups/TNRS.backup-remake` to back up TNRS before making changes to it. this provides a more recent restore point than the last import in case the changes mess things up. (however, the last import's backup is usually sufficient unless TNRS has been run since then.)
9993	06/20/2013 05:53 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): added VACUUM ANALYZE and runtime (50 s)
9992	06/20/2013 05:42 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): updated runtime (16 min)
9991	06/20/2013 05:09 PM	Aaron Marcuse-Kubitza	schemas/VegBIEN/taxonomy/higherPlantGroup.xlsx.src.txt: added Brad's comment that there are some holes in the Embryophyte subclasses list, and we need to validate it
9990	06/20/2013 04:49 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs: documented that when changing this table's schema, you must also make the same changes on vegbiendev. included sample util.set_col_types() call with runtime (4 min).
9989	06/20/2013 03:58 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): updated runtime (16 min)
9988	06/20/2013 03:32 PM	Aaron Marcuse-Kubitza	bugfix: inputs/.TNRS/schema.sql: tnrs_populate_fields(): need to schema-qualify invoked functions
9987	06/20/2013 03:29 PM	Aaron Marcuse-Kubitza	bugfix: inputs/.TNRS/schema.sql: tnrs_populate_fields(): Is_homonym: use the *_is_homonym flag for whichever of genus or family (in that order) is NOT NULL, rather than horizontal-ORing potentially NULL values together
9986	06/20/2013 03:22 PM	Aaron Marcuse-Kubitza	bugfix: inputs/.TNRS/schema.sql: family_is_homonym(), genus_is_homonym(): need to return NULL instead of false when input family/genus is NULL. EXISTS does not support this, so STRICT is used to provide this functionality automatically.
9985	06/20/2013 03:19 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: added family_is_homonym(), genus_is_homonym() and use them in tnrs_populate_fields()
9984	06/20/2013 03:15 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: score_ok(): changed to IMMUTABLE and STRICT
9983	06/20/2013 03:14 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): updated runtime (16 min)
9982	06/20/2013 02:41 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): never_homonym: use Author_score threshold to exclude matches that are too fuzzy to confirm the presence of a plant name author
9981	06/20/2013 02:38 PM	Aaron Marcuse-Kubitza	bugfix: inputs/.TNRS/schema.sql: tnrs_populate_fields(): *_is_homonym: also need to check that there was no Author_matched (i.e. that it could be a homonym). Is_homonym: use new never_homonym var.
9980	06/20/2013 02:18 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): updated runtime (18 min)
9979	06/20/2013 02:17 PM	Aaron Marcuse-Kubitza	inputs/import.stats.xls: Updated import times
9978	06/20/2013 02:07 PM	Aaron Marcuse-Kubitza	planning/workflow/bien3_architecture.pptx: stage II: removed Step prefix before stage #, which the other slides don't have
9977	06/20/2013 01:57 PM	Aaron Marcuse-Kubitza	added planning/workflow/bien3_architecture/stages.png
9976	06/20/2013 01:36 PM	Aaron Marcuse-Kubitza	added planning/workflow/bien3_architecture/stage_*.png
9975	06/20/2013 12:54 PM	Aaron Marcuse-Kubitza	added planning/workflow/bien3_architecture.pptx
9974	06/20/2013 08:20 AM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): when changing this function: UPDATE statement: include TNRS schema since it may not be in the search_path
9973	06/20/2013 08:14 AM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): Is_plant: also consider homonyms using new family_is_homonym, genus_is_homonym (see wiki.vegpath.org/Result_filtering#taxon_is_plant)
9972	06/20/2013 08:03 AM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs: added Is_homonym derived col (uses IRMNG.family_homonym_epithet, genus_homonym_epithet)
9971	06/20/2013 07:54 AM	Aaron Marcuse-Kubitza	schemas/vegbien.sql: re-ran `make schemas/public/reinstall; make schemas/remake` cycle, which apparently changed sort order of statements
9970	06/20/2013 07:11 AM	Aaron Marcuse-Kubitza	/README.TXT: Full database import: disk space check: updated minimum (to 300GB) for new import schema size. note that most of the space (166GB) is indexes, and even of the 87GB of data, only 20GB is from GBIF and 15GB from FIA (so most of it is duplication).
9969	06/20/2013 07:07 AM	Aaron Marcuse-Kubitza	added inputs/IRMNG/_homonym_epithet/map.csv, etc. (created by /run)
9968	06/20/2013 07:01 AM	Aaron Marcuse-Kubitza	bugfix: inputs/input.Makefile: `%/install %/header.csv: %/create.sql`: in noclobber mode, mark %/header.csv as .PRECIOUS so the existing file won't be deleted if the table already exists (causing an error exit)
9967	06/20/2013 06:54 AM	Aaron Marcuse-Kubitza	bugfix: lib/runscripts/table.run: remake_VegBIEN_mappings(): run yes using piped_cmd() so the SIGPIPE doesn't cause an errexit
9966	06/20/2013 06:45 AM	Aaron Marcuse-Kubitza	added inputs/IRMNG/{genus_homonym_epithet,family_homonym_epithet}/run, which inherit from ../table.run so that load_data() (which runs create.sql) is invoked
9965	06/20/2013 06:44 AM	Aaron Marcuse-Kubitza	added inputs/IRMNG/species_homonyms/new_terms.csv
9964	06/20/2013 06:43 AM	Aaron Marcuse-Kubitza	bugfix: added no-op inputs/IRMNG/Source/run so inputs/IRMNG/run would have something to invoke for it
9963	06/20/2013 06:35 AM	Aaron Marcuse-Kubitza	inputs/IRMNG/run: use lib/runscripts/datasrc_dir.run, which now provides import() and $subdirs
9962	06/20/2013 06:34 AM	Aaron Marcuse-Kubitza	lib/runscripts/datasrc_dir.run: extend import.run and provide an import() implementation that runs all the runscripts for import_order.txt subdirs
9961	06/20/2013 06:20 AM	Aaron Marcuse-Kubitza	lib/csvs.py: sniff(): support single-column spreadsheets by defaulting to the Excel dialect when the delimiter can't be determined
9960	06/20/2013 06:09 AM	Aaron Marcuse-Kubitza	inputs/IRMNG/: added family_homonym_epithet, genus_homonym_epithet lookup tables, which use util.all_same() to filter out internal Plantae homonyms
9959	06/20/2013 05:44 AM	Aaron Marcuse-Kubitza	schemas/util.sql: added all_same() aggregate
9958	06/20/2013 05:31 AM	Aaron Marcuse-Kubitza	schemas/util.sql: added not_empty(anyarray)
9957	06/20/2013 12:14 AM	Aaron Marcuse-Kubitza	schemas/util.sql: added not_null() (usable as an aggregate's FINALFUNC)
9956	06/20/2013 12:13 AM	Aaron Marcuse-Kubitza	schemas/util.sql: added not_null() (usable as an aggregate's FINALFUNC)
9955	06/19/2013 10:18 PM	Aaron Marcuse-Kubitza	bugfix: inputs/IRMNG/import_order.txt: need to specify order so that Source is first
9954	06/19/2013 10:16 PM	Aaron Marcuse-Kubitza	bugfix: inputs/IRMNG/*/map.csv: remapped Authority to scientificNameAuthorship instead of authors (now data_authors <VegCore.vegpath.org?data_authors> for clarity)
9953	06/19/2013 09:59 PM	Aaron Marcuse-Kubitza	inputs/IRMNG/map.csv: updated to scrubbed output names from */map.csv (/map.csv does not currently get scrubbed)
9952	06/19/2013 09:51 PM	Aaron Marcuse-Kubitza	bugfix: inputs/IRMNG/species_homonyms/header.csv, map.csv: reset input columns to DSV (delim-separated values) header. they had gotten changed to the output names in running map.csv with remake=1, causing it to be remade from the (renamed) staging tables.
9951	06/19/2013 08:54 PM	Aaron Marcuse-Kubitza	inputs/input.Makefile: $(_svnFilesGlob): added *Makefile
9950	06/19/2013 08:51 PM	Aaron Marcuse-Kubitza	/README.TXT: `make inputs/{upload,download}`: first run with test=1 to see what the diffs will be
9949	06/19/2013 08:47 PM	Aaron Marcuse-Kubitza	added inputs/IRMNG/, including runscripts to download the names. this is now the 2nd datasource after GBIF to use runscripts, and the 3rd after FIA/GBIF to use new-style import.
9948	06/19/2013 08:45 PM	Aaron Marcuse-Kubitza	inputs/input.Makefile: $(_svnFilesGlob): added *run (runscripts)
9947	06/19/2013 08:24 PM	Aaron Marcuse-Kubitza	lib/runscripts/table.run: import(): also run remake_VegBIEN_mappings() to accept the test output. this function was previously unused, but was left in for future use when lib/import.sh was translated to lib/runscripts/table.run (it was used in its import.sh form in inputs/FIA/occurrence_all/import).
9946	06/19/2013 08:21 PM	Aaron Marcuse-Kubitza	bugfix: lib/runscripts/table.run: remake_VegBIEN_mappings(): need to change to $top_dir before running `rm header.csv map.csv`
9945	06/19/2013 08:18 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: added in_top_dir()
9944	06/19/2013 08:00 PM	Aaron Marcuse-Kubitza	lib/runscripts/table.run: remake_VegBIEN_mappings(): only remake header.csv, map.csv if this target is being run directly, to avoid needing to remake them every time. for tables that are views, this instead requires them to be explicitly remade when the view columns change.
9943	06/19/2013 07:07 PM	Aaron Marcuse-Kubitza	bugfix: lib/runscripts/subdir.run: subdir_make(): only remake if $remake has been explicitly propagated to subdir_make() by using self_make
9942	06/19/2013 06:51 PM	Aaron Marcuse-Kubitza	lib/sh/make.sh: added deferred_check_target_exists alias and use it in check_fake_target_exists
9941	06/19/2013 06:41 PM	Aaron Marcuse-Kubitza	added lib/sh/web.sh with curl wrapper
9940	06/19/2013 06:41 PM	Aaron Marcuse-Kubitza	lib/sh/make.sh: added check_wildcard_target_exists alias
9939	06/19/2013 06:37 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: added wildcard1 alias

Project

General

Profile