Large Calculations#
Here, a calculation is considered large if it cannot be run, i.e. if it runs out of memory, it fails with strange errors or it just takes too long to complete.
There are various reasons why a calculation can be too large. 90% of the times it is because the user is making some mistakes and is trying to run a calculation larger than really needed. In the remaining 10% of the times, the calculation is genuinely large and the solution is to buy a larger machine, or to ask the OpenQuake engine developers to optimize the engine for the specific calculation that is giving issues.
The first things to do when you have a large calculation is to run the command oq info --report job.ini, that will
tell you essential information to estimate the size of the full calculation, in particular the number of hazard sites,
the number of ruptures, the number of assets and the most relevant parameters you are using. If generating the report
is slow, it means that there is something wrong with your calculation and you will never be able to run it completely
unless you reduce it.
The single most important parameter in the report is the number of effective ruptures, i.e. the number of ruptures after distance and magnitude filtering. For instance your report could contain numbers like the following:
#eff_ruptures 239,556
#tot_ruptures 8,454,592
This is an example of a computation which is potentially large - there are over 8 million ruptures generated by the model - but that in practice will be very fast, since 97% of the ruptures will be filtered away. The report gives a conservative estimate, in reality even more ruptures will be discarded.
It is very common to have an unhappy combinations of parameters in the job.ini file, like discretization parameters
that are too small. Then your source model with contain millions and millions of ruptures and the computation will
become impossibly slow or it will run out of memory. By playing with the parameters and producing various reports,
one can get an idea of how much a calculation can be reduced even before running it.
It is a good idea to read the section about Common mistakes.
Running large hazard calculations, especially ones with large logic trees, is an art, and there are various techniques that can be used to reduce an impossible calculation to a feasible one.
Reducing a calculation#
The first thing to do when you have a large calculation is to reduce it so that it can run in a reasonable amount of time. For instance you could reduce the number of sites, by considering a small geographic portion of the region interested, of by increasing the grid spacing. Once the calculation has been reduced, you can run it and determine what are the factors dominating the run time.
As we discussed in section about common mistakes, you may want to tweak the quadratic parameters (maximum_distance,
area_source_discretization, rupture_mesh_spacing, complex_fault_mesh_spacing). Also, you may want to choose
different GMPEs, since some are faster than others. You may want to play with the logic tree, to reduce the number of
realizations: this is especially important, in particular for event based calculation were the number of generated
ground motion fields is linear with the number of realizations.
Once you have tuned the reduced computation, you can have an idea of the time required for the full calculation. It will be less than linear with the number of sites, so if you reduced your sites by a factor of 100, the full computation will take a lot less than 100 times the time of the reduced calculation (fortunately). Still, the full calculation can be impossible because of the memory/data transfer requirements, especially in the case of event based calculations. Sometimes it is necessary to reduce your expectations. The examples below will discuss a few concrete cases. But first of all, we must stress an important point:
Our experience tells us that THE PERFORMANCE BOTTLENECKS OF THE
REDUCED CALCULATION ARE OFTEN DIFFERENT FROM THE BOTTLENECKS OF
THE FULL CALCULATION. Do not trust your performance intuition.
Collapsing the GMPE logic tree#
Some hazard models have GMPE logic trees which are insanely large. For instance the GMPE logic tree for the latest European model (ESHM20) contains 961,875 realizations. This causes two issues:
it is impossible to run a calculation with full enumeration, so one must use sampling
when one tries to increase the number of samples to study the stability of the mean hazard curves, the calculation runs out of memory
Fortunately, it is possible to compute the exact mean hazard curves by collapsing the GMPE logic tree. This is a simple as listing the name of the branchsets in the GMPE logic tree that one wants to collapse. For instance in the case of ESHM20 model there are the following 6 branchsets:
Shallow_Def (19 branches)
CratonModel (15 branches)
BCHydroSubIF (15 branches)
BCHydroSubIS (15 branches)
BCHydroSubVrancea (15 branches)
Volcanic (1 branch)
By setting in the job.ini the following parameters:
number_of_logic_tree_samples = 0
collapse_gsim_logic_tree = Shallow_Def CratonModel BCHydroSubIF BCHydroSubIS BCHydroSubVrancea Volcanic
it is possible to collapse completely the GMPE logic tree, i.e. going from 961,875 realizations to 1. Then the memory issues are solved and one can assess the correct values of the mean hazard curves. Then it is possible to compare with the value produce with sampling and assess how much they can be trusted.
NB: the collapse_gsim_logic_tree feature is rather old but only for engine versions >=3.13 it produces the exact
mean curves (using the AvgPoeGMPE); otherwise it will produce a different kind of collapsing (using the AvgGMPE).