Large Pandas Dataframe parallel processing

后端未结

关注

 2  2047

感情败类 2021-02-07 20:38

I am accessing a very large Pandas dataframe as a global variable. This variable is accessed in parallel via joblib.

Eg.

df = db.query(\"select id, a_lo


      
      
        
          2条回答        

        
                    
            
            
                         
                
              
              
                
                   礼貌的吻别
                                             
                
                
                (楼主)
            
              
              
                2021-02-07 21:21
              

            
            
                        
The entire DataFrame needs to be pickled and unpickled for each process created by joblib. In practice, this is very slow and also requires many times the memory of each.  

One solution is to store your data in HDF (df.to_hdf) using the table format.  You can then use select to select subsets of data for further processing.  In practice this will be too slow for interactive use.  It is also very complex, and your workers will need to store their work so that it can be consolidated in the final step. 

An alternative would be to explore numba.vectorize with target='parallel'.  This would require the use of NumPy arrays not Pandas objects, so it also has some complexity costs.

In the long run, dask is hoped to bring parallel execution to Pandas, but this is not something to expect soon.
    
             
                                                        
            
            
              
                
                0
              
                   
                
               讨论(0)
              
                                                  
              
              
                          
             
       
          
              
                                       
     查看其它2个回答


            
                         
                    


               
            
    发布评论:
    
         
                        
    
    提交评论 
  
  

                    
                    
                    
                        
                        
                         加载中...
                        
                    
                
          
                              			
        
        
        
          
            
            
              
              
            
    


                                 
              
            
                          
    

        
         
                验证码
                
                  
                
                
                   看不清?
                
              
                                  
                    
   
                 
             
              提交回复