I'm curious... between Whereby and Jitsi and I assume other browser-based video solutions relying on WebRTC...
...how big is the barrier these days to building a "videoconferencing platform" supporting millions of people... that runs on a single server?
Because if you need to do is build a pretty website that essentially just keeps track of meeting names and the names and IP addresses of participants...
...while each client is P2P-streaming their full-res videostream while speaking or other participants have them pinned... and every other client is P2P-streaming a low-res videostream to power the thumbnails (and similar decisions about which computer is the main audio source and when, or picking a single peer to serve as the audio mixer)...
What else is there to do, really?
(I mean obviously there's fancy stuff you can add like screensharing, chat, authentication, etc... and browser-specific bugfixes and quirks presumably...)
But are we at a point where anyone can write a functional videoconferencing platform in a week, and platforms are differentiating mainly on nicer UX and extra features?
Or is there something huge I'm missing here, where implementing WebRTC is somehow a lot harder than it seems, and/or still requires server farms to route the streams through in certain cases?
It doesn't scale well beyond a handful of people. You need *N bandwidth to send and receive, and without a thing called 'simulcast' (creating multiple, different quality stream simultaneously) which doesn't have good browser support the quality is defined by the lowest common denominator. A central server solves many, many issues that result in better quality.
Jitsi itself barely works on Firefox and not at all on mobile devices (without their app).
It actually works fine on Android browsers (checked in FF and Chrome), you just have to load the page in Desktop mode so that it stops pushing the apps to you.
Hopefully they'll reconsider their decision if they want it to get popular...
I rolled https://video.etherpad.org out within 5 minutes. It's a single command once Etherpad is installed (npm install ep_webrtc).
There is one complication most people don't realize -- Failed Reverse NAT traversal: For this you need a TURN server (I'm intentionally ignoring STUN for obvious reasons).
TURN servers have to route the actual media (video / audio) from user a <> b <> c but only if the user(s) can't directly connect. We hit Tb's a day through our TURN server and it gets expensive.
But complexity wise, it's an absolute doddle! Give it a go, if you have nodejs installed 90% of your work is done!
Silly question, because I tried to run Nextcloud Talk and ran into odd connection issues for a user who I believe is behind a corporate firewall and so I needed to stand up “coturn”: what’s the obvious reason for avoiding STUN? And what would you recommend as the simplest/best TURN server implementation?
Hosting cost is the biggest barrier to building a video conferencing solution which scales to millions of users. We setup a Jitsi meet instance and with just 6 parties it pegged a core on the server CPU at 50%.
Admittedly one of the users was on Firefox which causes CPU load to spike with Jitsi but either way video conferencing is bandwidth and processor intensive.
Otherwise the WebRTC technology is stable and works well across browsers - especially for audio. Just scaling it and getting folks to pay for it so it’s economically feasible to host is another thing.
Bandwidth and processing power are limiting factors. Our department tried to run a large Big Blue Button instance to support a dozen of conferences at the same time, ranging from 10-150 participants, all day long.
The experience says: you need hardware (not virtual), starting from 32 cores and 64GB RAM, it turned out it was not enough, added another machine, then another machine, ... and more.
...how big is the barrier these days to building a "videoconferencing platform" supporting millions of people... that runs on a single server?
Because if you need to do is build a pretty website that essentially just keeps track of meeting names and the names and IP addresses of participants...
...while each client is P2P-streaming their full-res videostream while speaking or other participants have them pinned... and every other client is P2P-streaming a low-res videostream to power the thumbnails (and similar decisions about which computer is the main audio source and when, or picking a single peer to serve as the audio mixer)...
What else is there to do, really?
(I mean obviously there's fancy stuff you can add like screensharing, chat, authentication, etc... and browser-specific bugfixes and quirks presumably...)
But are we at a point where anyone can write a functional videoconferencing platform in a week, and platforms are differentiating mainly on nicer UX and extra features?
Or is there something huge I'm missing here, where implementing WebRTC is somehow a lot harder than it seems, and/or still requires server farms to route the streams through in certain cases?