Changelog
All notable changes to the oda_data package are documented here.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[2.7.0] - 2026-06-15
This release refreshes the DAC1 indicator catalogue to match the current
OECD DAC1 flow-type classification. A number of aid types were reassigned to
different flow types upstream, which changes their indicator codes (the
DAC1.<flowtype>.<aidtype> middle segment). The underlying data is unchanged —
only the codes used to address it. If you reference any of the codes below
directly, update them.
This also fixes a bug that made the DAC1 indicator generator unrunnable
(dac1_aid_flow_type_mapping() read a non-existent flow_type key instead of
flowtype_code), and adds test coverage for the previously-untested generator.
Changed
- 17 DAC1 indicator codes changed (same data, new flow-type segment):
| Old code | New code | Indicator |
|---|---|---|
DAC1.50.5 |
DAC1.5.5 |
Official and private flows |
DAC1.37.415 |
DAC1.30.415 |
(reassigned to Private development finance) |
DAC1.50.420 |
DAC1.30.420 |
(reassigned to Private development finance) |
DAC1.50.425 |
DAC1.30.425 |
(reassigned to Private development finance) |
DAC1.50.3300 |
DAC1.37.3300 |
(reassigned to Other private market) |
DAC1.50.3320 |
DAC1.37.3320 |
(reassigned to Other private market) |
DAC1.50.3530 |
DAC1.37.3530 |
(reassigned to Other private market) |
DAC1.50.359 |
DAC1.37.359 |
(reassigned to Other private market) |
DAC1.50.3840 |
DAC1.37.3840 |
(reassigned to Other private market) |
DAC1.50.3860 |
DAC1.37.3860 |
(reassigned to Other private market) |
DAC1.50.3890 |
DAC1.37.3890 |
(reassigned to Other private market) |
DAC1.50.7530 |
DAC1.37.7530 |
(reassigned to Other private market) |
DAC1.10.11002 |
DAC1.40.11002 |
(reassigned to Non flow) |
DAC1.50.2231 |
DAC1.1021.2231 |
Memo: development finance in blended finance packages |
DAC1.50.2232 |
DAC1.1021.2232 |
Memo: … through funds and facilities |
DAC1.50.2233 |
DAC1.1021.2233 |
Memo: amounts mobilised from the private sector |
DAC1.50.2234 |
DAC1.1021.2234 |
Memo: … of which through guarantees |
- Two flow types added to the DAC1 flow-type vocabulary:
5("Total official and private flows") and1021("Memo items (mobilisation and blended finance)").
Added
- Two new DAC1 indicators:
DAC1.10.1623(Debt buybacks) andDAC1.37.1030(Offsetting entry for debt relief — private claims, principal). - Test coverage for the DAC1 indicator generator and its mapping loaders.
Fixed
dac1_aid_flow_type_mapping()raisedKeyError(readflow_typeinstead offlowtype_code), which broke DAC1 indicator regeneration.
[2.6.0] - 2026-04-28
This release reorganises how oda_data stores downloaded data on disk.
Caches now live in a standard per-user location instead of inside your
project folder, and a new oda_data.cache.* namespace gives you a clear,
typed way to inspect and manage them. Existing caches are migrated
automatically on first run. See Cache Management for the
full walkthrough.
Added
-
Per-user cache by default. Downloads now live under your OS's standard cache directory (
~/Library/Caches/oda-data/on macOS,~/.cache/oda-data/on Linux,%LOCALAPPDATA%\oda-data\on Windows), versioned per release. This means cache is shared across projects instead of duplicated in every.raw_data/folder, and old versions don't silently get reused after an upgrade. -
Override the cache location with
oda_data.set_cache_root(path)or theODA_DATA_CACHE_DIRenvironment variable — useful for shared volumes or CI runners with limited home-directory space. -
One place to manage the cache:
oda_data.cache. Inspect what's cached, clear specific parts, or invalidate a single dataset without touching the rest:
from oda_data import cache, CRSData
cache.size() # bytes per scope
cache.clear("raw") # drop only raw OECD zips
cache.invalidate(CRSData) # forget cached CRS, keep everything else
-
Per-call
refresh=Trueon every dataset'sread()to force a fresh download for one call without permanently clearing the cache. -
Automatic recovery from corrupt downloads. When a freshly downloaded zip fails its integrity check, the bad file is removed and the download is retried once. If both attempts fail, you get a clear
BulkPayloadCorrupterror pointing at the file and the reason.
Changed
set_data_path()no longer controls the cache. It now only sets where the package writes parquet exports (data you explicitly save out). If you used it to redirect cache storage, switch toset_cache_root()orODA_DATA_CACHE_DIR— the package will print a one-time deprecation warning to remind you. The old call still works through 2.x and is removed in 3.0.clear_cache(),enable_cache(), anddisable_cache()still work unchanged. They now delegate to the newcache.*namespace, so existing scripts keep running without edits.
Migration
On the first cache-touching call after upgrading, the package looks for
pre-2.6 caches in their old locations (./.raw_data/, the per-OS
oda-reader directory) and moves them into the new layout. Caches on
synced drives (Dropbox, iCloud, OneDrive) are skipped with a clear log
message — re-run with oda_data.cache.migrate(force=True) to override.
[2.5.1] - 2026-04-28
Fixed
- OECD CRS bulk downloads using the newer PKZIP Deflate64 compression
method no longer fail with
BadZipFile: File is not a zip file. Pulled in viaoda-reader >= 1.5.1. If you saw this error before upgrading, delete any stale file under youroda-readerbulk cache (path varies by OS) before retrying.
[2.5.0] - 2026-04-09
Changed
- Romania (code 77) is now classified as a DAC member/country (previously non-DAC), reflecting its new status as a DAC associate.
[2.4.2] - 2026-02-13
Fixed
- Memory cache returning stale results when a cached DataFrame lacks columns requested by a subsequent query.
[2.4.1] - 2025-12-19
Added
- CRS column mappings for
donor→provider_nameandrecipient→recipient_name.
[2.4.0] - 2025-12-19
Added
- New
DATA_TYPE_CODEfield to ODASchema and CRS column mapping for datatype_code column support.
Changed
- DAC2A bulk downloads now use dedicated
bulk_download_dac2a()function from oda-reader for improved reliability. - Measure filters are now skipped for DAC2A when using bulk downloads (consistent with CRS behavior).
- Updated oda-reader dependency from
>=1.3.1to>=1.4.1.
[2.3.2] - 2025-12-19
Added
- New Development Bank (code 1044) to provider groupings, CRS names, and DAC2A names.
- Eurasian Fund for Stabilization and Development (code 1041) to provider groupings, CRS names, and DAC2A names.
- UN Economic and Social Commission for Western Asia (code 1403) to provider groupings and DAC2A names.
[2.3.1] - 2025-12-15
Added
- European Investment Bank (EIB, code 919) to provider/donor mappings across DAC1, DAC2A, CRS, and provider groupings.
- New unspecified regional recipient codes: Southern Asia (6790), Micronesia (8600), Middle Africa (10280), Melanesia (10330), Polynesia (10350).
- Broad sector categories for top-level aggregation: Education, Health, Energy, General Environment Protection, Agriculture and Forestry & Fishing.
- New sector/purpose mappings including Conflict/Peace/Security (152), Trade Policies (331), Refugees (930), Humanitarian Aid (700).
Fixed
- Type conversion for code columns (
sector_code,purpose_code,donor_code,agency_code) when adding name columns to handle mixed types. - Typo in sector name: "Unallocated/ Unspecified" → "Unallocated/ Unspecified".
- Capitalization in broad sector groups: "government & Civil Society" → "Government & Civil Society".
- Sector mapping now uses fallback for unmapped sectors and fills missing values with "Unallocated/ Unspecified".
[2.3.0] - 2025-10-16
Added
- Comprehensive test suite with unit and integration tests
clean_parquet_file_in_batches()function for memory-efficient processing of large files- Thread-safe memory caching with
ThreadSafeMemoryCache - Manifest-based bulk cache tracking system
- Query cache manager for filtered dataset results
- Contributing guidelines and pre-commit hooks
Changed
- Complete caching refactor with three-tier architecture (memory, bulk, query caches)
- Improved thread and process safety using FileLock for cache coordination
- Better memory management with configurable cache size limits
- Atomic file operations for cache writes to prevent corruption
- Enhanced error handling in query filter construction
Fixed
- Cache corruption issues in multi-threaded/multi-process environments
- Memory issues when processing large bulk files (now processes in batches)
- Race conditions in cache initialization across threads
- Stale cache detection and automatic refresh logic
[2.2.2] - 2025-09-26
Fixed
- Bug with marker calculations
[2.2.1] - 2025-09-26
Added
- Access to sector imputations via
from oda_data import sector_imputations
Fixed
- Issues with filter passing given schema changes in bulk files on the OECD side
[2.1.2] - 2025-09-01
Fixed
- Caching paths now respect user-defined data directories and default to a
.raw_datafolder relative to the working directory
[2.1.1] - 2025-07-23
Fixed
- Bug where GNI may not get converted to constant prices even if a base year is specified
[2.1.0] - 2025-06-16
Changed
- Improved AidDataData to behave more like other Sources
[2.0.6] - 2025-06-16
Fixed
- Bug when trying to calculate multilateral imputations in constant prices
[2.0.5] - 2025-06-13
Fixed
- Bug caused by the Providers multisystem dataset using a form of pascal case
[2.0.4] - 2025-06-13
Changed
- CRS research indicators now use bulk downloads by default
[2.0.3] - 2025-06-13
Changed
- Better bulk file memory management
[2.0.2] - 2025-05-28
Added
- Functionality to calculate the official ODA/GNI
[2.0.1] - 2025-04-25
Changed
- Improved caching performance by keeping both memory and disk cache of parquet files
[2.0.0] - 2025-04-22
This major release is a complete refactoring of the oda-data package. It is now faster,
more stable, and better organized.
Changed
- Complete package refactoring with improved performance and stability
- BREAKING: Major API changes - please refer to the Migration Guide for details
- Versions ~1.5.x will remain supported until at least August 2025 to allow time to migrate workflows
Version 1.x Releases
[1.5.0] - 2024-11-29
Changed
- Updated requirements to pydeflate >=2.0
- Removed climate indicators (given methodological challenges inherent in OECD data). For access to climate data, please see the climate-finance package
[1.4.3] - 2024-11-29
Fixed
- JSON validation error for recipient groupings
[1.4.2] - 2024-11-26
Fixed
- Donors and recipient groupings to fully align with recent schemas
[1.4.1] - 2024-10-11
Fixed
- Bug with how certain files are stored, moving them from feather to parquet
[1.4.0] - 2024-10-11
This release introduces significant changes to how raw data files are managed. It is strongly recommended that all users update to this version.
Changed
- Default storage format changed from feather to parquet files, allowing oda_data to leverage predicate pushdown and more efficiently load only the data it needs
- Removed data download tools from oda_data in favor of using the tools via oda-reader
- oda-reader package now uses the new data-explorer API and bulk downloads instead of relying on the old (and now inaccessible) bulk download service
[1.3.3] - 2024-09-16
Fixed
- Issues reading bulk files from the OECD (given that the bulk download service no longer exists)
[1.3.1] - 2024-07-16
Fixed
- Schema of the temporary fix to align with the expected CRS schema from the bulk download service
[1.3.0] - 2024-07-16
Added
- Workaround for the OECD bulk download service, which is down following the release of the new OECD website
- Uses a full CRS file shared by the OECD (note: nearly 1GB and can take a long time to download on slow connections)
[1.2.0] - 2024-04-05
Changed
- Now uses
oda_readerto download data for DAC1 and DAC2a directly from the API - Data is converted to the .Stat schema to ensure full backwards compatibility
- Updated dependencies
Deprecated
- .Stat schema will be deprecated in a future version in favor of the explorer API schema
[1.1.6] - 2024-03-14
Changed
- Updated pydeflate dependency to deal with data download issue
[1.1.5] - 2024-03-07
Fixed
- Bug introduced by changes in the OECD bulk download service
[1.1.4] - 2024-03-01
Fixed
- Constant non-USD currencies bug for imputed sectors calculations
[1.1.3] - 2024-02-29
Fixed
- Sorting bug (arrow)
[1.1.2] - 2024-02-29
Added
- Support for reading the CRS from 1973-2004
Fixed
- Removed a warning on pandas stack (for future behavior)
[1.1.1] - 2024-02-29
Security
- Security updates to dependencies
[1.1.0] - 2024-02-29
Added
- New indicators to separately produce multilateral sector spending shares and imputed multilateral spending totals
- Improved, automated method to map multilateral CRS spending (by agency) to the multilateral "channels" used in the multisystem database
- Tools to group purpose codes following ONE's sector groupings
[1.0.11] - 2024-01-04
Fixed
- Key COVID indicators
[1.0.10] - 2023-12-11
Changed
- Added UTF8 encoding
[1.0.7] - 2023-12-11
Security
- Updated requirements for security
[1.0.6] - 2023-10-21
Fixed
- Bug caused by new readme files in the bulk download service file
[1.0.5] - 2023-08-24
Changed
- Updated how the CRS codes are fetched given the connection issues outlined in the notes for 1.0.4
- Updated how the indicators that use the
multisystemdatabase work - the OECD quietly changed the output format of the database, which broke the parsing of the data. The new format is now supported
[1.0.4] - 2023-08-24
Added
- Backup solution to download bulk files from the OECD website using
selenium(given an insecure SSL certificate that causes the normal download usingrequeststo fail) - Dependencies:
seleniumandwebdriver-manager
[1.0.3] - 2023-06-12
Changed
- Updated requirements (pydeflate) to address the same OECD data bug as in 1.0.2
[1.0.2] - 2023-06-12
Fixed
- Encoding bug that affected CRS data given a new file encoding from the OECD bulk downloads
Changed
- Updated requirements
[1.0.1] - 2023-04-13
Changed
- Updated requirements to a newer version of pydeflate, given data quality issues with the latest OECD release
[1.0.0] - 2023-02-20
First major release of oda_data. We have settled on the basic functionality of the package and the basic API.
Changed
- Updated requirements
Early Releases (v0.x)
[0.4.1] - 2023-01-30
Changed
- Updated requirements
[0.4.0] - 2023-01-30
Added
- Indicators for climate finance data
[0.3.5] - 2023-01-12
Fixed
- Issues with research indicators in non-USD data
[0.3.4] - 2023-01-12
Fixed
- Issues with gender data
[0.3.3] - 2023-01-13
Fixed
- Issues with multilateral non core ODA
[0.3.2] - 2023-01-12
Fixed
- Issues with multilateral sector imputations
[0.3.0] - 2023-01-10
Added
- ONE Core ODA indicators (flows, ge, linked ge), including 'non Core' indicators
- "Official definition" total ODA indicator
[0.2.5] - 2022-12-21
Added
- Ability to retrieve COVID-19 indicators
[0.2.3] - 2022-12-16
Fixed
- ODA GNI indicators, which returned mostly invalid data from the source
- Typo in the ODA GNI indicator name
- How
OECDClientdeals with adding shares to indicators for which shares don't make sense
[0.2.1] - 2022-12-16
Changed
- Download data for indicator automatically if not available in data folder
[0.2.0] - 2022-12-16
Added
- Method to OECDClient to add a "share" column to the output data
- Method to OECDClient to add a "gni_share" column to the output data
Changed
OECDClient().load_indicator()now accepts a list of indicators as input
[0.1.10] - 2022-12-09
Added
- Total (ODA + OOF, excluding export credits) indicator for the CRS
[0.1.9] - 2022-12-07
Added
- Ability to request a 'one_linked' indicator - these indicators are composed of a main indicator which is completed by a fallback indicator when values are missing
- Option to get a simplified/summarized dataframe by calling
.simplify_output_df()on theOECDClientobject - Documentation for the
OECDClientclass
Changed
- How indicators are grouped when requesting a 'one' indicator
[0.1.8] - 2022-11-29
Added
- More comprehensive tests of all core functionalities
- Tool to extract CRS codes from the DAC CRS code list
[0.1.7] - 2022-11-24
Fixed
- Issue with trying to set a file path for both oda_data and pydeflate
[0.1.0] - 2022-11-24
First release of oda_data