Showing posts with label class. Show all posts
Showing posts with label class. Show all posts

Sunday, March 12, 2017

On Why Quitting is Terrible Failing

Cooper's Hawk just down the trail. Just using a point-and-shoot camera.
Many failures may precede success, but failing within an institution is just failing.


I am trying to suss out what is an acceptable level of failure. I grew up with the idea that failing at anything was bad: school, piano, that stupid cake-walk back in cub scouts, and many other venues. I’ll just say that I was so upset at the cake-walk they gave me some consolation pie. But back to the point: failure was unacceptable and very upsetting.


School was and is the prime place of failure being unacceptable. In some ways I envy the kids that don’t care when they fail, but in many others I don’t, since I would be very surprised if they are more successful than myself. Though I am certainly not a paragon of success, just very comfortable.


What does failing lead to within school as an institution? Retaking a grade at the worst save dropping out. I never failed in primary or secondary with my worst grade a C. This kept me from having a free ride through UW. The only fear of failure ever was on specific tasks, and really only on public performances. But also the schools were small enough and loose enough that there were definitely ways of working through failure that a teacher would consider equivalent.


College was much tougher for me. Fortunately I had skin in the game, money, debts, but it still was a very tough thing for me to finish. I had failures and retakes and then I graduated. Did I learn more? Yes, but as much as I probably should have? No. I would have done better in that case.


But that brings me to my master’s degree. I really like the computer science program that I can take online very far away from Georgia Tech. I will highly recommend the program to anyone looking at getting a CS MS. However, it should come with a warning: Gatech will punish you for failing.


Wait a second, of course they should, getting a bad grade because you didn’t understand the material or didn’t do the work or didn’t regurgitate the correct answer on a closed book test, directly leads to punishment, right? Sure, I have bad time management, better than before, still bad, but it’s not quite what I mean: Gatech will punish you for doing better a second time. This is much more controversial.


In the startup world there is a saying of “fail fast” that is, try many different things to see what sticks. In a good company failure is celebrated as a learning opportunity. In normal companies failure is grounds for firing. Not showing up to work? Obvious fireable offense, barring reasons. Trying an interesting project that flops? Should never be punished except when the repercussions can’t be absorbed or mitigated.


But in the world of academics it is a whole different animal. If you fail it goes on your record and that is it, always a mark against you. In some cases a permanent mark against you can be a good warning flag for employers or future institutions: “This person has done very poorly in this area.” But more often it is much more of an indicator of environment and what the person is going through at the time they get the grade. For those persistent folks out there who have perfect time management, I applaud you.


In my undergrad degree  the university had a policy of replacing the first grade with the second if it was better. Then after three attempts they started to average the grades. There was room for big improvements, but then it was tempered to make sure someone didn’t just take advantage of the system. I understand Georgia Tech’s system might want the grad students to be able to discern if they are going to pass the class, but this is nearly impossible when the professors don’t have a good idea of how the grades might be “bumped” or curved.


Either way I have dropped both classes. Failure is tough, freeing, but I can’t get over the fact that I spent money and vast amounts of time on these classes. Fortunately for Computer Vision I am pretty sure that I will be able to use/improve the code from the first five assignments that I did. I am not sure I will take databases again. I learned things, but I don’t want to learn it in a formal way, rather just apply what I need to do my job.


I will probably sign up for one summer course, one fall course, and so on. Very annoying, but at least it won’t kill me in the process.


It does free up time: travel, bird class, fencing, bird watching, archery. So many things plus more writing and actually adjusting and using my 3d-printer.

Failing sucks, and quitting is failing, it just feels worse.

Saturday, February 04, 2017

End of Week 5 2017

Already a month into 2017. A glorious future where everyone is happy, or: Everyone is quite happy to be unhappy if given the choice. I really don’t know what to say that won’t enrage you, I would really love to say that no one has anything to worry about from the new way things are, but that isn’t true. However, it is really hard to believe anything, even if my initial reaction is rage my immediate follow-up is skepticism.

That is the way democracy falls, or even republics if you want to avoid anything with the letters “dem” in it. The more we can be convinced to be unreasonably outraged or not care is a win. The first is easy to crackdown on, and the second just lets them do whatever they want. Try to take a second when something strikes you as outrageous. Let the sweet feeling of indignation die down and try to get some corroboration rather than just believing anything you see.

In less divisive ideas: I am nearly a month into two courses. It is quite a bit of work, but I am enjoying my Computer Vision course and I am quite happy with the team I have for Databases. For the vision course I wish I had known all of the information that I know now and what I am about to learn last year at about this time, it would have really helped in the robot competition.

The biggest problem with the classes is free time. If I take too much free time, such as writing this, I am seriously procrastinating. Or even worse is finding a game like Avorion. It is a really well done little game (scales a galaxy), even though it is still in early access. But I should really just quit playing any games until the break, otherwise they will break me. That is a tough thing to write. Maybe in the future I will write about the games I play, though I suppose it might be boring. Just a quick search shows that there are already many people who call themselves “The Boring Gamer.” Weird.

If I can manage these two classes and get through summer and two more in the fall, I can graduate! Then I will have sooo much free time it won’t be funny. Then I can:
  • Write more
  • Join more things like archery and orchestra
  • Play games
  • Hang out with my wife without making her feeling guilty about keeping me from homework
  • Bird watch more
  • Travel
  • Run out of time to do things.
Yeah, that sounds great.

Saturday, October 04, 2014

To Take or Not to Take the Subway



Introduction to Data Science Capstone
Udacity
The New York Subway system spans 842 miles of track with 468 stations and transports about 5 million people a day. But those are just large numbers, first we need to get a feel for the data. What does the precipitation look like throughout the year? What does the ridership look like on an average day? Can we predict ridership with the given information?

The dataset I am using to answer these questions is turnstile data from 2013 collected over the entire city, however, I am limiting mine to the most consistent station for reporting data: Wall Street. I also signed up for an API on Weather Underground to get hourly data for the same year. Let’s take a look at precipitation throughout 2013.
Precipitation in New York 2013: Shows rain and snow amounts throughout the year.

This graph has all 365 days of the year with the precipitation of each hour arrayed along the horizontal axis. The alpha of the dots shows us the intensity of the rain in the hour. The most intense rain in an hour recorded 1.06 inches. Given that information and that the alpha setting needs to be between 0 and 1, I normalized all the rest of the rain data against that value giving us an alpha gradient and a decent visualization.

It gives us a sense of the precipitation, but we also need a good view of ridership throughout the year. However just creating a line graph or a scatter plot won’t really help us visualize the data. The first thing to try would be averaging all the hours to get an average day’s ridership.

After some file size issues, data had glitches from when a turnstile would cut-out, and I also found that the reporting time of every four hours to be problematic, especially when a single turnstile would report during an off hour.


I found that the only station to get nearly continuous updates was the Wall Street station. It is only missing a total of 21 reports for the entire year. So I ran it through a series of filters designed to get rid of steps and glitches and came up with a graph that looks much more like a counting line:
Running it through some SQL queries in pandas I got the average ridership per hour so that we can compare it to our rain data, at least visually.



The interesting thing with the entries data is that morning rush hour is quite a bit more than the really early morning, but it never really slacks off at lunchtime, in fact there are more and more entries until a peak of about 2,750 entries at 18:00.

But does this correspond at all to our weather data? As it has one similar axis I could flip it so that the axes match up and then have a dual axis along the bottom.


This is the result with the red line as the average day/workday throughout the year and the orange line denoting the weekend days. It might just be a fluke of the climate, but it seems June’s heavy rain-showers correspond quite nicely with the uptick in traffic at 18:00. However, this is also about the time most people are leaving the non-residential district to head back home after a day of, ironically, pushing numbers to get better results.

Although visually biasing the graph really doesn't tell us what we want to know, and this is where we turn to regression.


It seems that the best predictor of subway foot traffic is time. With over a thousand more hours predicted within 1000 commuters compared to using just weather, which is more than ten percent of the data set, deciding when to ride the subway will affect the number of people riding with you more than if it is raining.

However, I was wondering what a different algorithm might be able to tell us. What might a machine learning model be able to predict, possibly trained on different parts of the year? So I turned to scikit-learn to go through a few different algorithms.

At first I tried a Bayesian Ridge which is very similar to the algorithm used in the class and that I modified to use on this data. Its results were nearly the same, although I was able to get some very interesting overfitting using the entirety of the year as a learning set. Otherwise it seemed that somewhere in April had a set of data that was the best at predicting the entire year. When I checked the r-squared value I was a bit surprised: 0.23. Not as good as I was hoping despite the graphs showing much of the predictions within one thousand passengers.
AdaBoost compares quite favorably to the regression used in class. It also has an r-squared value of nearly 0.53.

Both algorithms have problems predicting the evening rush-hour numbers, however AdaBoost seems to recover for later in the night.
I also tried the AdaBoost algorithm which sets up multiple prediction agents rather than just weight the inputs. With this algorithm the predictions got an r-squared of 0.52, quite a bit better. Still the driving factor behind all of these predictions is time, and I could probably get an even better result by separating the weekends and holidays out of the weekdays. And a problem with the AdaBoost algorithm: it isn’t completely repeatable with minimum controls as it gives a different distribution each time, but seems to maintain the r-squared value.
Difference of ridership against prediction (red) and precipitation (blue) show that not the prediction algorithm was not greatly affected by precipitation, getting predictions nearly right and quite wrong. The deepest red dots are predictions off by nearly twice the highest average traffic, about 6000 riders.

Generally it seems that the majority of people using the Wall Street station use it all of the time, that and the people who don’t either continue walking or using other transportation despite most weather.

The last thing to consider is the different mapreduce options. First of all the whole dataset isn’t truly that large, at 60MB it certainly bogged down my computer and ran into system timeouts, but removing those limits and putting in a few more Gigs of RAM to deal with copies easily would go a long way to solve these problems, not really in the realm of true need of MapReduce.

To get into this realm we might look at something that could be a bit more complicated. Real time tracking of entrances and exits on the subway system to get a much finer grain picture of the entire system would certainly get close to qualifying, especially if you wanted to more than just swim in the data.

Another future application of real time systems maybe routing traffic with a pay-or-get-paid system for taking city-wide data and destinations and crunching it down to route traffic. If a person needs to be somewhere quickly they can pay a fee to take a faster route while those taking slower routes would get paid to wait. With the advent of self-driving cars this might not be as onerous as one might think as the people stuck in the slow traffic could possibly work, that or their company can pay for them to use the faster routes to be in more quickly.

But dealing with hundreds of thousands of inputs every minute would definitely need a system that could split up the tasks between servers or possibly server farms. Being able to predict problems or heavy loads becomes a necessary issue.

Most of these traffic issues, whether subway or cars, can be predicted most easily using times and events, the only time that weather is truly going to affect traffic patterns will be in extreme or catastrophic situations. If you are pinning your hopes on getting a seat on the subway because of weather patterns then you might get a seat just before the subway shuts down during a snow day.

References:


https://www.udacity.com/course/ud359 - Udacity course
http://ggplot2.org/ - python plotting library
http://www.wunderground.com/weather/api/ - Weather data
http://web.mta.info/developers/download.html - New York subway data
http://scikit-learn.org/stable/ - Regression algorithms

Wednesday, March 07, 2012

Chugging Along

Getting back into the swing of things.These last few weeks have been pretty horrendous for fencing and running.

Many things conspired to limit fencing, plans and cancellations, but finally got to fence last night. I fenced a bit and practiced some too. After being up the mountain yesterday, and getting up to do 1.5 miles in the morning I certainly wasn't at top performance, however it was definitely good to fence. I certainly need to improve, as we have a person that can soundly beat me. It is quite amazing how much of a motivator that is, in fact it will be an excellent learning opportunity.

I really need to work on my point control, and in that direction I have made a ball with a string through it. Really simple, but attaching a clamp to the beam and then dangling the ball from it should help my point control become better. Or rather I will improve my point control. Another thing is reactions, I need to definitely put in place stock actions as well as leave the door open for flexibility. I see things just fine, but as I get more tired I make bigger actions, which are most decidedly not restful. They also cause pieces of blades to break off and go flying across the gym. So some shopping needs to be done to find a few dry blades, and a glove as mine is finally wearing out, Leon Paul certainly makes great gloves.

I ran 1.5 miles yesterday and 2 miles this morning. I haven't run in 2 weeks mostly because of the rain. I dislike running in the rain, but I made a deal with myself Monday night, I would run at least once a week whether or not it was raining in the morning, Monday being the day. That would allow for my shoes to dry out. But it was actually nice Tuesday morning, so I didn't need to get wet. When I reach 300 miles in my shoes I think they will be relegated as I get a new pair, but I am still about 110 miles from that goal. An old pair of shoes, or just even a second pair that can get wet without worrying about tomorrow.

I also didn't do too well on the most recent self-driving car homework. It wasn't so much programming as putting together matrices for a Kalman Filter. The fact that I didn't have to take linear algebra in college probably doesn't help. Also waiting until the last moment doesn't help either. This week I started in on the lectures so I can hopefully get that done well before this weekend and start in on the homework. It is focused on Particle Filters, supposedly a much easier idea. I sure hope so, but I suspect that we will have to know all three including a histogram discrete filter. It should continue to be interesting, and I am learning python at least.

Tuesday, February 28, 2012

Update, These Days...

Doing quite a bit better these days so far as motivation goes. Although really need to be more disciplined for exercise, but I really dislike running in the rain. I guess I could get another set of shoes so that I can let them dry, but I also need to remember to take out the insoles and flatten them so that I don't get blisters from slightly curled insoles.

Fencing is probably completely off for this week. First Mauna Loa School is doing something so we can't fence there, and Thursday I am going up the mountain and may be up there until 21:00 which would be at least an hour after the longest winded fencers are still there. Oh well. Also in fencing news I won't be going to any interesting tournaments when I go to the ESC Conference in March for work. And I looked at May when Jessie and I will be attending her sister's graduation, nothing, but maybe we are just too far out. There are tournaments happening, but they are all qualifiers for nationals and since I am not from their area, I cannot be a part.

The ESC Conference is in the last part of March, San Jose, CA. I will mostly be learning about and summarizing technologies I see for work, but I will also be looking and just seeing what I can see. It will certainly be interesting. I suppose that I will be pretty busy but this is going to feel quite weird going to a conference like this. I have really only heard of and read about industry conferences... I did attend a week long course in a conference like thing in Boston 2 summers ago, but that was not just a huge conglomeration of companies and industry leaders. It was experts teaching about what they knew, quite a bit different.

Ah, just looked up fencing in San Jose... The hotel I am staying at is about a block from The Fencing Center, so could have pretty interesting evenings, or at least on Monday and Wednesday. It seems that I will need to bring gear... I should double-check my connections. Gear is definitely unwieldy to travel with, although it would be checked and have plenty of room for my other clothing.

It will definitely be an interesting month.

Featured Post

Allergy

John studied himself in the mirror as best he could through tears. Red, puffy eyes stared back at him, a running nose already leaked just a ...