Well, if you look at how we've done things over the years:
- We used to have our own machines in our own data centers
- Then we started renting machines in data centers
- We then moved to the cloud model where we would get compute capacity on demand. But the minimum unit was an hour
- But what if you could deploy your code and you were only charged for the compute and memory you take for the fulfillment of that request? That is what the "serverless" model is. You don't constantly run a server like apache; instead, when you receive a user Request, the relevant function is called, executed and results returned to the user. You are billed for the ram ⨉ CPU.
This has many benefits:
- For low traffic sites, this has significant cost savings
- For high traffic sites, this is auto-scaling without thinking about launching machine instances
This isn't without problems, naturally. This lends itself well to certain type of problems better than others. For example, if you need lots of hot-cache data, the response times on a serverless stack would be slower.
But as you can see from the above, this is the logical direction cloud computing will evolve. Infrastructure will truly be shared and you should be able to extract efficiencies down to the minute.
I wonder if that's really true (that it scales better than running more-or-less bog standard CGI across a fleet of thousands of servers, behind a load balancer).
If I've understood lambda correctly, rather than spinning up a process (as with CGI), it spins up an entire vm/container. I suppose it might do the fast-cgi thing - spin up a container, and keep it running while there are requests coming in, and then kill it off.
I'm sure there are other benefits of "container-on-demand" vs "process-on-demand" -- but I'm not sure "scales better" is one of them. Well, I'm sure it "scales better" in the sense of organization (human resources), but not necessarily in terms of machine resources.
The concept goes back even further. It's really transaction processing as used in the IBM Customer Information Control System, first used in 1964. Load small program image on demand, run it, discard it. (Or, optionally, reuse it for another transaction.) This is usually tied to a database, and if the transaction fails, the database changes are rolled back. This is how IBM mainframes do transactions. 48 years later, versions of CICS are still in wide active use and are supported IBM products.
We used to have "web hotels". Then they grew server side programming support, such as Perl and later PHP. Customers were at first not separated at all but later they were with Virtuozzo and similar systems.
This was later rebranded PaaS, to differentiate from cheap PHP hosting of yore. Some has now rebranded to serverless. It's not so much a "direction" as it is market differentiation.
GAE will start up a machine occasionally to deal with increased demand and keep them on for 15 minutes at least. With something like Lambda (my old favourite was PiCloud) I can scale out to 50 processes for 10 seconds and only pay for 500 seconds of processing time.
For me, most of the benefit comes from data processing, where my usage is very bursty. The classic example used on most of the tutorials is image thumbnails/scaling/similar. All you want to say is "Do X to every file in this folder of 50k images" and let someone else handle starting and stopping as many machines as they can in order to get this done as quickly as possible.
No, heroku and to a lesser extant appengine are running at the whole application level. Lambda and google's, and azure's functions are literally individual functions that are deployed separately. They can be piped together to make a whole app but don't need to be.
- We used to have our own machines in our own data centers
- Then we started renting machines in data centers
- We then moved to the cloud model where we would get compute capacity on demand. But the minimum unit was an hour
- But what if you could deploy your code and you were only charged for the compute and memory you take for the fulfillment of that request? That is what the "serverless" model is. You don't constantly run a server like apache; instead, when you receive a user Request, the relevant function is called, executed and results returned to the user. You are billed for the ram ⨉ CPU.
This has many benefits:
- For low traffic sites, this has significant cost savings
- For high traffic sites, this is auto-scaling without thinking about launching machine instances
This isn't without problems, naturally. This lends itself well to certain type of problems better than others. For example, if you need lots of hot-cache data, the response times on a serverless stack would be slower.
But as you can see from the above, this is the logical direction cloud computing will evolve. Infrastructure will truly be shared and you should be able to extract efficiencies down to the minute.