I took the example “Out-of-Core Training on MNIST” from
The only code element I have modified is
results =
NetTrain[lenet, trainingDataFiles, All,
ValidationSet -> testDataFiles, MaxTrainingRounds -> 3,
TargetDevice -> "GPU"]
So, I added 'TargetDevice → “GPU” ’ which is typically not a problem. But in this example I get the following message:
The Stack Trace for NetTrain::interr2
The issue does not occur with “CPU” as TargetDevice. How to solve the problem?
jeromel
(Jérôme Louradour)
December 3, 2019, 2:06pm
2
What is the content of Internal`$LastInternalFailure after the failure?
Are you running a 12.0?
Yes, I am running V12 on Windows 10.
I do not know what the “content of Internal`$LastInternalFailure after the failure” is. How to get it?
jeromel
(Jérôme Louradour)
December 3, 2019, 3:00pm
4
Just run the command that fails.
And after that, evaluate:
Internal`$LastInternalFailure
(as suggested in the error message)
Then copy paste the output here
After a re-start of Mathematica the problem does not occur again.
My assumption is that Clear[“Global`*”] on top of the notebook does not clear the GPU memory and this was the reason for the failure.
What is the best command to start this kind of notebook (re-)evaluations with a clean memory of all needed devices?
jeromel
(Jérôme Louradour)
December 3, 2019, 5:12pm
6
Clear should not have any effect on the GPU memory.
My idea when suggesting to reboot was more the following:
When putting a computer in a sleep mode and using it again, the GPU might be in a weird state and WL may have problems to connect to it. I wanted to make sure that was not the case.
No, the computer did not went in a sleep mode during my activities.
jeromel
(Jérôme Louradour)
December 5, 2019, 11:45am
8
OK. The error in case of weird GPU state after sleep mode should anyway be:
At least you don’t have any error now.
If it happens again, check the value of Internal`$LastInternalFailure and you can report here. Thanks!