I'm currently calculating the hypergeometric survival function (1-cdf) for a large number of comparisons (~900M). As it stands, the scipy implementation is relatively slow (at least compared to the R version, which is ~100x faster):
def calculate_hypergeometric_overlap_nums(object_ontology_count, overlaps):
'''
Calculate hypergeometric 1-cdf
'''
fftot = 40071
ffb = 375
if overlaps <= 0: return [1.0]
overlaps_adj = overlaps - 1
val = stats.hypergeom.sf(overlaps_adj, fftot, object_ontology_count, ffb)
return [val]
results = list(tqdm(map(self.calculate_hypergeometric_overlap_nums, object_ontology_counts, oe_overlaps)))
> 1394400it [02:32, 9156.94it/s]
The above is the timing for ~1.4M comparisons. I see the following SO post that discusses modifying some underlying C libraries:
How to efficiently vectorise hypergeomentric CDF calculation?
Which seem to speed up the process by ~100x, but this process is a bit over my head. Are there other ways of quickly calculating scipy.stats.hypergeom.sf ?