Saturday, October 4, 2008
Google Video gave it the old college try.
A good analysis of how ecosystem monitoring will make your life easier
Sunday, September 21, 2008
How Google's outages hurt your business
Let's skip over the lack of transparency from Google during the event, except to say that it's pretty sad that it took so long to at least admit to the problem. To their defense, they claim it affected a small number of clients. And Google is generally open about their problems, so we'll give them a pass on this one.
Why does it matter to you?
Unlike downtime at GMail and Google Reader, a SaaS services like Custom Search being down is a big deal. Why? Because if you were using Custom Search, to your visitors it looks like YOU are down. Imagine being a customer of Smug Mug, visiting their help page and ending up with a really slow or broken search. Would you blame Google or Smug Mug? Sure many customers would probably blame themselves, but just the possibility that your perceived uptime and user experience is dependent on a third party (that you have no control over or insight into) should give you pause. How are you supposed to even know that these services are down? Imagine if it was something more critical to your business like your ad network or the payment processing system?
Is SaaS doomed?
In spite of these dangers, the benefit of using SaaS solutions is very strong. Why bother building and hosting something outside your core competency when a service out there does it for you. You can read about the benefits of SaaS here, here, here, here, and here. I doubt I have to convince you of that. So the question is how you can continue to reap the benefits of SaaS while minimizing your exposure to problems you can't control. Is there a solution?
A solution
The key to a successful SaaS implementation is having real time access to the uptime and performance of the SaaS solutions your business relies on. If you knew right away that Google's Custom Search solution was down, at the least you could react put up a friendly message for your visitors ("Don't blame us, it's Google's fault!"). Even better you'd have a fail-over plan in place to switch to another solution. Same thing if this was an ad network or a payment system that went down. You would have some control over your user's experiences, and would no longer have to pray that all of your solution providers are up 100% of the time (good luck!). Without this knowledge, you're either assuming these services never go down, or you don't realize that your visitors have no idea that the issues aren't your fault.
The company I work for recently launched a solution that deals with this very need. It's all about working together with your SaaS providers, sharing performance and uptime data, and being able to see the same data your providers are seeing. As with most problems, it often times boils down to opening up the communication lines.
As more businesses come to rely on SaaS solutions, the more exposure these business will have to this kind of "perceived" downtime. The naive solution is to expect 100% uptime. The real solution is to know when that downtime does occur, and to have a plan of action.
Nice to see some talk of transperency in the blogosphere
"Google isn't exactly known as the most transparent company in the world, but they're light years ahead of Apple - a company that in some ways they share a kinship with when it comes to their reputation for innovation. Apple (or for that matter any big company) can learn a lot about radical transparency, customer service and PR from Google, even though they're hardly perfect here."He goes on to review the various places that Google and Apple make public their bugs and known issues. What's missing here obviously is any mention of transparency in uptime and performance. But to fill in the gaps, as we've seen previously, Google does a much better job here as well.
Saturday, September 20, 2008
Robert Scoble hosting a webinar on scalability
I'm certainly registering. Will be interesting to get their thoughts on load testing, and how they plan for such large amounts of unpredictable load.Avoiding the Fail Whale - Thursday, October 9th 1pm EST
Building a server environment that’s scalable and reliable can be tough, especially when your traffic goes “nuts” virtually overnight. Fast Company Live presents a special one-hour live webinar, moderated by Robert Scoble and featuring a panel of tech leaders from companies big to small who are facing these very issues.
Confirmed guests include:
- Matt Mullenweg: Founder of Automattic, the company behind WordPress.
- Paul Bucheit: One of the founders of FriendFeed and the creator of Gmail.
- Nat Brown: CTO of iLike, a music community service that had one of the first Facebook apps.
The discussion will cover architectural choices, growth hurdles and how the panelists overcame them. The first half-hour will be devoted to the panel discussion, while the second half-hour will be open to live questions from registered webinar attendees.
Sunday, September 7, 2008
Gamer's Bill of Rights?
Twitter showing improved uptime
TechCrunch quotes co-founder Biz Stone:
What I like most is the details that Twitter provides on their blog describing where the issues stem from:"Twitter has been making great progress in terms of uptime and reliability. Fail Whale sightings are far less frequent these days thanks to our efforts but we still have a long journey ahead. Last month we saw 99.88% uptime and so far this month we are at 99.96%. Our engineering and operations teams have been taking a very methodical approach to improving Twitter. We’re using the word “craftsmanship” to characterize our work here at the office. Reliability and dependability continue to be top on or list of key goals."
"I've always respected a good sense of pacing. It's easy to be fast and loose, but it takes a certain discipline, foresight, and patience to guide something through the right way. For most of Twitter's early days, pacing could be considered an unattainable luxury. Our effort started with a bang and quickly accelerated to a disconcerting velocity that never let up. We found ourselves reacting to situations instead of crafting solutions and features we wanted to make.Kudos to Twitter for not only the improved uptime, but for keeping it's users in the loop on things that generally are discussed only behind closed doors.
With nearly two years at full speed, thousands of successes (with as many mistakes), and countless lessons learned, we've finally discovered our rhythm as a team. By carefully regrouping all aspects of our work, breaking the problem down into smaller parts, and iterating rapidly, Twitter, Inc. is poised to bring a new kind of communication to every part of the world."
