Google Colaboratory: misleading information about its GPU (only 5% RAM available to some users)

后端 未结 9 616
滥情空心
滥情空心 2020-12-02 03:40

update: this question is related to Google Colab\'s \"Notebook settings: Hardware accelerator: GPU\". This question was written before the \"TPU\" option was added.

相关标签:
9条回答
  • 2020-12-02 04:07

    Im not sure if this blacklisting is true! Its rather possible, that the cores are shared among users. I ran also the test, and my results are the following:

    Gen RAM Free: 12.9 GB  | Proc size: 142.8 MB
    GPU RAM Free: 11441MB | Used: 0MB | Util   0% | Total 11441MB
    

    It seems im getting also full core. However i ran it a few times, and i got the same result. Maybe i will repeat this check a few times during the day to see if there is any change.

    0 讨论(0)
  • 2020-12-02 04:08

    Last night I ran your snippet and got exactly what you got:

    Gen RAM Free: 11.6 GB  | Proc size: 666.0 MB
    GPU RAM Free: 566MB | Used: 10873MB | Util  95% | Total 11439MB
    

    but today:

    Gen RAM Free: 12.2 GB  I Proc size: 131.5 MB
    GPU RAM Free: 11439MB | Used: 0MB | Util   0% | Total 11439MB
    

    I think the most probable reason is the GPUs are shared among VMs, so each time you restart the runtime you have chance to switch the GPU, and there is also probability you switch to one that is being used by other users.

    UPDATED: It turns out that I can use GPU normally even when the GPU RAM Free is 504 MB, which I thought as the cause of ResourceExhaustedError I got last night.

    0 讨论(0)
  • 2020-12-02 04:11

    Find the Python3 pid and kill the pid. Please see the below image

    Note: kill only python3(pid=130) not jupyter python(122).

    0 讨论(0)
  • 2020-12-02 04:12

    So to prevent another dozen of answers suggesting invalid in the context of this thread suggestion to !kill -9 -1, let's close this thread:

    The answer is simple:

    As of this writing Google simply gives only 5% of GPU to some of us, whereas 100% to the others. Period.

    dec-2019 update: The problem still exists - this question's upvotes continue still.

    mar-2019 update: A year later a Google employee @AmiF commented on the state of things, stating that the problem doesn't exist, and anybody who seems to have this problem needs to simply reset their runtime to recover memory. Yet, the upvotes continue, which to me this tells that the problem still exists, despite @AmiF's suggestion to the contrary.

    dec-2018 update: I have a theory that Google may have a blacklist of certain accounts, or perhaps browser fingerprints, when its robots detect a non-standard behavior. It could be a total coincidence, but for quite some time I had an issue with Google Re-captcha on any website that happened to require it, where I'd have to go through dozens of puzzles before I'd be allowed through, often taking me 10+ min to accomplish. This lasted for many months. All of a sudden as of this month I get no puzzles at all and any google re-captcha gets resolved with just a single mouse click, as it used to be almost a year ago.

    And why I'm telling this story? Well, because at the same time I was given 100% of the GPU RAM on Colab. That's why my suspicion is that if you are on a theoretical Google black list then you aren't being trusted to be given a lot of resources for free. I wonder if any of you find the same correlation between the limited GPU access and the Re-captcha nightmare. As I said, it could be totally a coincidence as well.

    0 讨论(0)
  • 2020-12-02 04:12

    If you execute a cell that just has
    !kill -9 -1
    in it, that'll cause all of your runtime's state (including memory, filesystem, and GPU) to be wiped clean and restarted. Wait 30-60s and press the CONNECT button at the top-right to reconnect.

    0 讨论(0)
  • 2020-12-02 04:12

    Restart Jupyter IPython Kernel:

    !pkill -9 -f ipykernel_launcher
    
    0 讨论(0)
提交回复
热议问题