Posts

Day 23

Image
When I came in today, two of the models (the second and third one from yesterday) were still training, but the first one had stopped after the 20th epoch due to a “CUDA error: invalid device ordinal.” So I spent the morning working on my presentation and figuring out ways to change the code. The models didn’t finish until 3:30 pm, and it turns out that 50 epochs is not the way to go. Especially for the paired dataset one… not good. I tested the one that got stopped at 20 epochs and the outputs were decent (the color of the sky is closer), so I’m running that one with the unpaired dataset again for 20 epochs and also for 10 epochs. I’m also running it using the paired dataset for 20 epochs (not enough memory to do the 10 epoch one but that should be fast tomorrow). I’m not sure what the issue is with the paired image results, because they should be better than the unpaired, but it seems that the discriminator outperforms the generators early on, or there could be overfitting occurring...

Day 22

Image
When I came in this morning all three of the models were done training. After testing each of them and seeing the results, I had to abandon the simplified model since it seems to think the sky is red (among other problems). (Ahhhhhh)               The original model fared better. Here’s a result from using the unpaired dataset. The colors are pretty dull and there are weird outlines between the sky and the trees/buildings but it looks better at least. And here's a RGB to thermal translation. I realize I should probably include the ground truth for all of these. There are thousands of testing images, but I might cut it down so it’s easier for me to match up images for the presentation. As for using the paired dataset, the outputted images somehow look worse and the graphs from the training were also strange. Most of the images have white dots over them, and while there are warmer colors it lost the blue… I’m wondering if I need to ch...

Day 21

Last Friday my advisors showed me how to connect to the RIT servers from home so I worked on my project over the weekend. After deconstructing my training code and reconstructing it again I finally figured out it was the metric calculations that were causing the accumulating memory usage. After fixing it I was able to run the code once on the simplified model (however I didn’t get any graphs from it because of an error), but now that I'm sure it’s working I’m going to train the original model. The system administrator also taught me how to use a screen session so I can detach the computer and still have the code running in the background (in case the computer loses connection, which happened twice at home). Right now I’m training three models at the same time- it’s slower (it’s been almost 7 hours and the one with the simplified model has gone through around 15/30 epochs and the ones with the original model have gone through around 10/30 epochs) but I think it’ll still take less ti...

Day 20

I realized this morning that the problem might not just be with the lack of memory (foreshadowing...) so I spent the morning cutting down the amount of memory my model was using by simplifying the network layers and making the crop, batch size, etc. smaller. But as I found out later, I think the issue is with my training code rather than the model. At 1 pm I went to a master’s thesis defense by one of the students in the visual perception lab. Her presentation was very interesting and also engaged the audience. Later in the afternoon the system administrator finished installing the GPU, but when I tried to run my program it still used up too much memory, which was surprising and a little concerning. The memory usage is supposed to plateau after a certain point, but mine just keeps accumulating until it reaches the maximum. I'm looking into different reasons as to why that's happening, and I've gotten the accumulation to slow (but not stop yet). I was hoping to train ov...

Day 19

I mostly spent the day trying to find a place to train my model. First the system administrators set up a CIS account for me so I could access Grissom, but after many technical issues (computer wouldn’t let me use Grissom, screen froze, etc.) it turns out that Grissom doesn't have enough memory either because of all the people using it (but I appreciate their efforts). Ultimately, we decided the best solution was to install a new GPU into a machine the system administrators have lying around and using that to run my code. They said it should be done by tomorrow afternoon, so hopefully I can train over the weekend.  Today was also the Undergraduate Research Symposium. I went with Jocelyn during lunch and heard the keynote speaker, Jason Babcock, talk about his research and design of eye-trackers, his company, Positive Science, and his experiences at RIT. Later in the afternoon, we also went to look at the posters and listen to an oral presentation (I was intrigued because part of t...

Day 18

The desktop was supposed to be repaired today, but no one showed up to fix it. Since I can't run the program or train my model I couldn’t really do that much today, but I worked on the code as best I could by separating different parts and running “simulations” to make sure they would work (eg. graphs, etc.). The code for printing out images should also work, but I’m not sure because there might be an issue with the output tensor dimensions from the model. Finally, I added a PSNR calculation into the testing code. I also started looking at the pix2pix code that I’ll be comparing my code with, although that might not be part of the final presentation. I went to a PhD dissertation today (titled Ultrafast Laser Polishing for Optical Fabrication). The research was interesting and was presented really well, so that even a person who knew nothing about the topic (me) could follow along.

Day 17

Today my advisors got a more detailed description of the error in the model code, which indicated that there was an issue with the size and crop of the images. After changing them to multiples of 256 (because of the way I defined the U-Net structure), it now works! The next roadblock was the computer memory, which I guess can no longer handle running the code. I tried running my training code this morning with my model, but it was really slow and was killed by the time I got back from lunch and the master's thesis defense I went to. I ran it again, but even after two hours it was still on the first epoch, and eventually it also ended because the CPU didn’t have enough memory. I’m using the CPU because the GPU doesn’t have enough memory either (even less because other people might be using it) and the computer in the lab is still broken. I emailed my advisor though and I think we'll work out an alternative tomorrow. I also found out that training can take 2-3 days, which is lon...