Speed up scipy.stats.hypergeom calculations

Viewed 109

I'm currently calculating the hypergeometric survival function (1-cdf) for a large number of comparisons (~900M). As it stands, the scipy implementation is relatively slow (at least compared to the R version, which is ~100x faster):

    
    def calculate_hypergeometric_overlap_nums(object_ontology_count, overlaps):
        '''
        Calculate hypergeometric 1-cdf 
        '''
        fftot = 40071
        ffb = 375
        

        if overlaps <= 0: return [1.0]
        overlaps_adj = overlaps - 1

        val = stats.hypergeom.sf(overlaps_adj, fftot, object_ontology_count, ffb)
    
        return [val]
       


results = list(tqdm(map(self.calculate_hypergeometric_overlap_nums, object_ontology_counts, oe_overlaps)))

 > 1394400it [02:32, 9156.94it/s] 

The above is the timing for ~1.4M comparisons. I see the following SO post that discusses modifying some underlying C libraries:

How to efficiently vectorise hypergeomentric CDF calculation?

Which seem to speed up the process by ~100x, but this process is a bit over my head. Are there other ways of quickly calculating scipy.stats.hypergeom.sf ?

0 Answers
Related