/ - Changes - BIEN 3 - NCEAS Projects

root @ 9938

#	Date	Author	Comment
9938	06/19/2013 06:37 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: added echo1()
9937	06/19/2013 06:32 PM	Aaron Marcuse-Kubitza	lib/runscripts/table.run: load_data(): first make sure schema is installed
9936	06/19/2013 06:31 PM	Aaron Marcuse-Kubitza	lib/runscripts/table.run: added datasrc_make_install()
9935	06/19/2013 06:31 PM	Aaron Marcuse-Kubitza	table_make_install(): take $install_log as an overridable kw param to support install logs in different locations
9934	06/19/2013 05:55 PM	Aaron Marcuse-Kubitza	lib/runscripts/table.run: load_data(): split noclobber functionality into separate table_make_install() function, which can be used by other install-related targets
9933	06/19/2013 03:27 PM	Aaron Marcuse-Kubitza	added schemas/VegBIEN/taxonomy/higherPlantGroup.xlsx.src.txt with Brad's description of how the names were chosen
9932	06/19/2013 02:37 PM	Aaron Marcuse-Kubitza	added schemas/VegBIEN/taxonomy/higherPlantGroup.xlsx
9931	06/19/2013 01:50 PM	Aaron Marcuse-Kubitza	schemas/VegBIEN/planning/taxonomy/: moved non-VegBIEN-specific resources to planning/resources/taxonomy/. this includes Brad's all-important Nomenclature_excerpt.ppt with the Latin taxonomic hierarchy suffixes on slide 5.
9930	06/19/2013 11:02 AM	Aaron Marcuse-Kubitza	bugfix: schemas/vegbien.sql: taxon_trait_view: use the TNRS-scrubbed name from ScrubbedTaxon when available
9929	06/19/2013 10:36 AM	Aaron Marcuse-Kubitza	schemas/vegbien.sql: split geoscrub_input_view's new-row-only filtering into separate view geoscrub_input_new, so that the full geoscrub_input rows are still available. the reduction in geoscrub_input from eliminating the already-scrubbed rows was only 280,000 (5076500 - 4799173) out of a possible 1.7 million (1707970), so it makes sense to just run geoscrubbing on the full input. (the lower-than-expected reduction is most likely due to rows from pre-refresh data being present in the original geoscrub_output table, which have been replaced by different, post-refresh input rows.)
9928	06/19/2013 10:18 AM	Aaron Marcuse-Kubitza	added exports/_archive/
9927	06/19/2013 10:17 AM	Aaron Marcuse-Kubitza	mappings/VegCore-VegBIEN.csv: genus->taxonlabel.taxonomicname: use new _filter_genus() (see r9882)
9926	06/19/2013 10:15 AM	Aaron Marcuse-Kubitza	backups/TNRS.backup.md5: updated
9925	06/19/2013 09:43 AM	Aaron Marcuse-Kubitza	bin/make_analytical_db: use new mk_table() instead of TRUNCATE/INSERT
9924	06/19/2013 09:41 AM	Aaron Marcuse-Kubitza	bin/make_analytical_db: added mk_table() and use it in mk_analytical_table()
9923	06/19/2013 09:30 AM	Aaron Marcuse-Kubitza	schemas/vegbien.sql: higher_plant_group_nodes: ferns and allies: added Lycopodiophyta node, as requested by Brad in the conference call (wiki.vegpath.org/2013-06-13_conference_call)
9922	06/19/2013 09:30 AM	Aaron Marcuse-Kubitza	schemas/vegbien.sql: geoscrub_input_view: exclude rows that have already been geoscrubbed, by anti-joining on geoscrub_output
9921	06/19/2013 09:11 AM	Aaron Marcuse-Kubitza	inputs/.geoscrub/geoscrub_output/postprocess.sql: set decimallatitude, decimallongitude types to double precision to facilitate joining with other double precision values
9920	06/19/2013 09:02 AM	Aaron Marcuse-Kubitza	inputs/.geoscrub/geoscrub_output/postprocess.sql: coords index: added rest of input columns so this can be used to check the existence of a result by input. added runtime (55 s). use idempotent create_if_not_exists().
9919	06/19/2013 08:17 AM	Aaron Marcuse-Kubitza	schemas/vegbien.sql: higher_plant_group_nodes: ferns and allies: added Lycopodiophyta node, as requested by Brad in the conference call (wiki.vegpath.org/2013-06-13_conference_call)
9918	06/19/2013 08:08 AM	Aaron Marcuse-Kubitza	bugfix: schemas/vegbien.sql: higher_plant_group_nodes: removed ferns and allies nodes Anthocerotophyta, Marchantiophyta, Bryophyta, which were incorrectly said to be part of this clade in the BIEN2 analytical DB overview (/planning/workflow/validation/BIEN2_Analytical_DB_overview.docx > p. 13 bottom > last ¶). see http://wiki.vegpath.org/2013-06-13_conference_call#fix-higher_plant_group_nodes-mapping .
9917	06/19/2013 07:58 AM	Aaron Marcuse-Kubitza	bugfix: /Makefile: postgres-Linux: phpPgAdmin: added steps to configure it for Apache 2.4
9916	06/18/2013 07:49 PM	Aaron Marcuse-Kubitza	/run: geoscrub_input/make(): documented runtime (40 s)
9915	06/18/2013 06:22 PM	Aaron Marcuse-Kubitza	bin/make_analytical_db: added `/run export_` to make the geoscrub_input CSV export
9914	06/18/2013 06:21 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs_populate_fields(): removed no longer needed casts of *_score to double precision
9913	06/18/2013 06:06 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs: *_score: changed type to double precision because these fields are always floats. this also avoids the need to manually cast them to double precision each time they are used.
9912	06/18/2013 05:55 PM	Aaron Marcuse-Kubitza	lib/tnrs.py: HTTP requests: rewrapped lines
9911	06/18/2013 05:53 PM	Aaron Marcuse-Kubitza	lib/tnrs.py: updated HTTP requests to match current web app
9910	06/18/2013 05:51 PM	Aaron Marcuse-Kubitza	bugfix: lib/tnrs.py: download_request_template: changed dirty to true (to match the current web app), which is apparently needed to apply the source_sorting setting to the downloaded TSV in addition to the GUI results
9909	06/18/2013 05:29 PM	Aaron Marcuse-Kubitza	lib/tnrs.py: retrieval_request_template: turned source_sorting back off, because it causes any match from the first source to always be used, even if it has a lower match score than the match from the other source. (Brad confirms that this should be off.) I think we had this on originally to ensure that only Tropicos results were used when available, rather than USDA when it was a better match. * note that due to a bug in the web app, this change will not actually be effective, because the source_sorting option is only applied to the GUI results, not the downloaded TSV. *
9908	06/18/2013 04:27 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: tnrs: Name_number: changed type to integer so it would sort numerically
9907	06/18/2013 04:24 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: added pkey on Time_submitted, Name_number
9906	06/18/2013 04:21 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: changed Name_submitted pkey to a unique constraint to allow adding a pkey on Time_submitted, Name_number instead
9905	06/18/2013 04:14 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/schema.sql: Time_submitted, Name_number: added NOT NULL constraints so that they can be used in a unique constraint
9904	06/18/2013 02:21 PM	Aaron Marcuse-Kubitza	lib/tnrs.py: submission_request_template: include GCC in addition to Tropicos, because it provides more synonyms than Tropicos for Asteraceae, and the accepted names still match the Tropicos backbone (https://projects.nceas.ucsb.edu/nceas/projects/bien/wiki/2013-06-13_conference_call#include-GCC-when-running-TNRS)
9903	06/17/2013 08:18 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/tnrs/tnrs.make: removed no longer needed end time, now that the total runtime is printed
9902	06/17/2013 08:17 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/tnrs/tnrs.make: print the total runtime using `time`
9901	06/17/2013 08:14 PM	Aaron Marcuse-Kubitza	inputs/.TNRS/tnrs/tnrs.make: include the end time in addition to the start time so that the total runtime can be calculated
9900	06/17/2013 07:59 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: command-specific alternate stdin/stdout/stderr: choice of 40/41/42: added mnemonic that 4 looks like A for Alternate
9899	06/14/2013 07:42 AM	Aaron Marcuse-Kubitza	bugfix: /README.TXT: Full database import: added step to remove any leftover TNRS lockfile. usually, the PID in it would not exist, but sometimes it now refers to a different, active process which blocks tnrs.make.
9898	06/13/2013 01:00 PM	Aaron Marcuse-Kubitza	schemas/vegbien.sql: allow public_ to view lookup tables (cultivated_family_locations, higher_plant_group_nodes)
9897	06/13/2013 08:13 AM	Aaron Marcuse-Kubitza	added backups/TNRS.backup.md5, vegbien.r9459.backup.md5
9896	06/12/2013 02:45 PM	Aaron Marcuse-Kubitza	bugfix: lib/sh/local.sh: sync_upload(): need to use --no-group to prevent the group from being reset to aaronmk upon download from jupiter (which uses group aaronmk instead of bien). use ./fix_perms to set the group of all files to bien. also use --no-owner in case running as root.
9895	06/12/2013 02:12 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: removed sync_download(). use swap=1 sync_upload() instead.
9894	06/12/2013 02:11 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: removed download(). use swap=1 upload, or swap=1 upload_caller, instead. this avoids having separate upload()/download() pairs for every caller of upload(), because you can instead just set swap=1.
9893	06/12/2013 01:50 PM	Aaron Marcuse-Kubitza	bugfix: lib/sh/sync.sh: upload(): don't kw_params $swap because this unexports it, preventing put from seeing it. instead, use echo_vars to print it.
9892	06/12/2013 01:35 PM	Aaron Marcuse-Kubitza	added bin/sync_upload, a wrapper around sync_upload()
9891	06/12/2013 01:23 PM	Aaron Marcuse-Kubitza	added bin/sync_upload, a wrapper around sync_upload()
9890	06/12/2013 01:00 PM	Aaron Marcuse-Kubitza	bugfix: lib/sh/sync.sh: upload(): only add `--exclude="**"` if there are --includes. this enables running upload() without paths to upload all files.
9889	06/12/2013 12:56 PM	Aaron Marcuse-Kubitza	lib/sh/sync.sh: upload(): support passing -- options to put, which will not be run through the path->--include processing
9888	06/12/2013 12:55 PM	Aaron Marcuse-Kubitza	bugfix: lib/sh/sync.sh: upload(): added missing `local args=()` initializer
9887	06/12/2013 12:18 PM	Aaron Marcuse-Kubitza	/README.TXT: Full database import: On local machine: added step to do steps under Maintenance > "to synchronize vegbiendev, jupiter, and your local machine", which is needed in addition to `make inputs/upload` since that doesn't handle overwrites or deletions
9886	06/12/2013 12:10 PM	Aaron Marcuse-Kubitza	/README.TXT: Maintenance: to synchronize vegbiendev, jupiter, and your local machine: added warning that you should pay careful attention to all files that will be deleted or overwritten (as the three machines are often out of sync)
9885	06/12/2013 11:26 AM	Aaron Marcuse-Kubitza	added inputs/GBIF/_MySQL/GBIFPortalDB-2013-02-20.data.0.preamble.sql
9884	06/12/2013 11:17 AM	Aaron Marcuse-Kubitza	/README.TXT: Full database import: make inputs/{upload,download}: run them first with `test=1` to see what the changes will be
9883	06/12/2013 11:12 AM	Aaron Marcuse-Kubitza	/README.TXT: Full database import: `svn up`: use --force to avoid errors about existing files
9882	06/12/2013 10:49 AM	Aaron Marcuse-Kubitza	mappings/VegCore-VegBIEN.csv: genus->taxonlabel.taxonomicname: filter out genera that contain numbers (using new _filter_genus()), which break TNRS and prevent it from matching any other parts of the name. later, these genera can instead be moved to the end of the name, where TNRS will correctly match them as Unmatched_terms.
9881	06/12/2013 10:48 AM	Aaron Marcuse-Kubitza	bugfix: inputs/VegBIEN/: added _no_import to disable import for this folder, since this is actually just an entry in web/datasources/ with VegPath redirection links, rather than an input to the import process. this fixes "schema "VegBIEN" does not exist" errors generated in `make test`.
9880	06/12/2013 10:45 AM	Aaron Marcuse-Kubitza	inputs/input.Makefile: $(dontImport): also support putting a _no_import file at the top level in the datasource to exclude the entire datasource
9879	06/12/2013 10:30 AM	Aaron Marcuse-Kubitza	bugfix: lib/sh/local.sh: removed make() override, which is no longer needed now that its operations are performed by verbosity_compat(), and which caused errors by setting $verbosity to the invalid value ""
9878	06/12/2013 10:28 AM	Aaron Marcuse-Kubitza	bugfix: lib/sh/util.sh: verbosity_compat(): always use default verbosity (`unset verbosity`) when verbosity == 1, regardless of whether the caller has set $verbosity to the special value "", because $verbosity is supposed to be an integer field and "" is not supported by most functions that use $verbosity. in cases where a util.sh script is invoked, it will set $verbosity back to the default value 1, so this will function as before for util.sh scripts and fix $verbosity for scripts that use a different verbosity system.
9877	06/12/2013 10:05 AM	Aaron Marcuse-Kubitza	added inputs/GBIF/raw_occurrence_record_plants/table.tsv.md5
9876	06/12/2013 09:51 AM	Aaron Marcuse-Kubitza	inputs/GBIF/raw_occurrence_record_plants/test.xml.ref: regenerated. updated for new staging table input columns, which are now the same as the output columns.
9875	06/12/2013 09:41 AM	Aaron Marcuse-Kubitza	bugfix: inputs/input.Makefile: %/VegBIEN.csv: use header from map.csv instead of the new columns, so that source.shortname is set to GBIF instead of VegCore
9874	06/12/2013 09:24 AM	Aaron Marcuse-Kubitza	inputs/input.Makefile: %/VegBIEN.csv: when a runscript is available, instead map the output columns of map.csv to VegBIEN, because the columns have been renamed in the staging table
9873	06/12/2013 08:32 AM	Aaron Marcuse-Kubitza	inputs/GBIF/raw_occurrence_record_plants/VegBIEN.csv: regenerated, which adds row_num input col
9872	06/12/2013 08:16 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: echo_func(): check can_log at beginning of function, so that the resource-intensive func_loc (which calls `readlink -f`) does not need to be called if echo_cmd would not log anything at the current verbosity
9871	06/12/2013 08:02 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: echo_func(): removed no longer used $minor flag. use `clog++... echo_func` instead.
9870	06/12/2013 07:25 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: verbosity_compat(): don't make $verbosity a local var of this function, because then the changes will not be visible to the caller
9869	06/12/2013 07:12 AM	Aaron Marcuse-Kubitza	bugfix: bin/make: use verbosity_compat because some make-invoked commands (e.g. bin/map) don't support verbosity=""
9868	06/12/2013 07:10 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: command(): command__exec(): use verbosity_compat to support commands that don't support verbosity=""
9867	06/12/2013 07:09 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: added verbosity_compat(), for use with commands that don't support verbosity=""
9866	06/12/2013 06:47 AM	Aaron Marcuse-Kubitza	bugfix: lib/sh/local.sh: make(): when invoking overridden func, need make__make_sh
9865	06/12/2013 06:46 AM	Aaron Marcuse-Kubitza	bugfix: lib/sh/util.sh: self, self_sys aliases: need to remove any func_override suffix __* from the FUNCNAME
9864	06/12/2013 06:35 AM	Aaron Marcuse-Kubitza	bugfix: inputs/GBIF/import_order.txt, run: updated raw_occurrence_record/ to raw_occurrence_record_plants/
9863	06/12/2013 06:32 AM	Aaron Marcuse-Kubitza	inputs/FIA/occurrence_all/test.xml.ref: update inserted row count
9862	06/12/2013 06:27 AM	Aaron Marcuse-Kubitza	bugfix: bin/make: include local.sh so that its default verbosity-setting make() override will be used
9861	06/12/2013 06:26 AM	Aaron Marcuse-Kubitza	lib/sh/local.sh: added make() override, which uses the default verbosity (i.e. verbosity="") when verbosity == 1. scripts that use lib/sql.py (which uses $verbosity) have different default verbosities, and this default should not be overriden by an env var, unless a higher verbosity has been set.
9860	06/12/2013 06:24 AM	Aaron Marcuse-Kubitza	lib/sh/local.sh: added missing include of make.sh, used by root_make()
9859	06/12/2013 05:57 AM	Aaron Marcuse-Kubitza	schemas/vegbien.sql: added _filter_genus()
9858	06/12/2013 04:47 AM	Aaron Marcuse-Kubitza	inputs/GBIF/raw_occurrence_record_plants/run: import() runtime: specified that this does not include table.tsv.gz/make()
9857	06/12/2013 04:07 AM	Aaron Marcuse-Kubitza	inputs/GBIF/raw_occurrence_record_plants/postprocess.sql: Remove institutions that we have direct data for: # duplicates: added revision #
9856	06/12/2013 04:07 AM	Aaron Marcuse-Kubitza	inputs/GBIF/raw_occurrence_record_plants/postprocess.sql: Remove institutions that we have direct data for: documented that there are 4.5 million duplicates (59,998,354 rows before - 55,417,646 rows after = 4,580,708)
9855	06/12/2013 03:49 AM	Aaron Marcuse-Kubitza	inputs/GBIF/raw_occurrence_record_plants/postprocess.sql: Remove institutions that we have direct data for: added rerun time (~0 thanks to index, so no problem doing the DELETE each time postprocess.sql is run)
9854	06/12/2013 03:25 AM	Aaron Marcuse-Kubitza	*{.sh,run}: use simpler .rel() instead of `. "$(dirname "${BASH_SOURCE⁰}")"/...` for relative includes
9853	06/12/2013 03:24 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: added .rel()
9852	06/12/2013 02:52 AM	Aaron Marcuse-Kubitza	schemas/VegCore/VegCore.pg.sql: regenerated from installed schema. a Linux bug that caused constraints to be output in reverse order has now been fixed, as of 12.04.2 (broken in 12.04.1).
9851	06/12/2013 02:48 AM	Aaron Marcuse-Kubitza	bugfix: inputs/GBIF/_MySQL/MySQL_schema, MySQL_data: sed: put {} commands on their own line to work on Mac
9850	06/12/2013 02:11 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: .(): put echo_func on its own line for clarity
9849	06/12/2013 02:10 AM	Aaron Marcuse-Kubitza	lib/sh/util.sh: .(): added echo_func (2 log_levels up because it's internal)
9848	06/12/2013 02:06 AM	Aaron Marcuse-Kubitza	schemas/VegCore/mk_derived: use new lib/sh/local.sh instead of lib/import.sh (a precursor to util.sh, etc. still used by inputs/FIA/)
9847	06/11/2013 07:50 PM	Aaron Marcuse-Kubitza	bugfix: load_data(): verbosity_min: use verbosity_min='' so that csv2db's default verbosity (3) is used, instead of setting the verbosity directly to 3, which caused the log++ logging output from bin/make to be echoed at verbosity 3, creating cluttered output
9846	06/11/2013 07:43 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: verbosity_min(): support value '', which sets verbosity=''
9845	06/11/2013 06:40 PM	Aaron Marcuse-Kubitza	bugfix: inputs/GBIF/raw_occurrence_record_plants/postprocess.sql: updated column names to match the renamings in map.csv, which are now performed on the staging table itself
9844	06/11/2013 06:38 PM	Aaron Marcuse-Kubitza	lib/sh/util.sh: run_args_cmd(): time the command so that the runtime of the outer runscript target (i.e. the command run from the shell) is printed at the end of the output, like in bin/make
9843	06/11/2013 05:56 PM	Aaron Marcuse-Kubitza	bugfix: inputs/input.Makefile: %/install: don't run $(cleanup) if it has already been run by $(import_install_), so that it doesn't run twice
9842	06/11/2013 05:54 PM	Aaron Marcuse-Kubitza	inputs/input.Makefile: %/postprocess: don't run postprocess.sql if it is supposed to be run by a runscript, because postprocess.sql may then depend on additional steps the runscript runs before it
9841	06/11/2013 05:25 PM	Aaron Marcuse-Kubitza	lib/runscripts/table.run: import(): use self_make on load_data so that the remake status determines whether the table is reinstalled
9840	06/11/2013 05:23 PM	Aaron Marcuse-Kubitza	bugfix: lib/runscripts/mysql.table.run: import(): added missing set_make_vars, needed by self_make
9839	06/11/2013 05:21 PM	Aaron Marcuse-Kubitza	bugfix: lib/runscripts/table.run: load_data(): need to use $_remake instead of $remake when using set_make_vars

Project

General

Profile