October 20208 BEHIND THE DISRUPTION OF OPEN SOURCE TECHNOLOGYBy Guan Wang, Analytics Specialist, Digital and Smart Analytics, Swiss Re [SWX: SREN]Open source makes technology cheaper and betterIf you are a developer, you can't really work around open source technologies. Your Operating System can be Linux. You can have MySQL or MongoDB as your database. From Hadoop, Spark to Pandas, from PyCharm to Jupyter, from sklearn to Tensorflow and PyTorch, open sourced software has been everywhere in your development process. Not only software, but also open source machine learning models have been extensively used in production systems, like those computer vision pretrained models generated from ImageNet, or Natural Language Processing pretrained models like word embeddings or even BERT. When we are talking about how AI is changing the world, we are really talking about how Open Source is changing the world. Most state-of-art tools and algorithms are published as academic research paper and open sourced as repositories on Gihub, and continuously adopted and improved by developers all over the world.Chinese companies were long considered as only imitators. Now companies like Baidu, Alibaba and Tencent are becoming strong power in the open source community. Open sourcing their own technology can promote themselves to attract developers to join their ecosystem and build products upon it. Sometimes contribution to a core open source program like Tensorflow is also a good testimony and advertisement to prove the technology capabilities. Nowadays not only internet companies but also traditional enterprises start to have their own open source policy.With open source, even a junior student can start using the most advanced technology within a fairly short time. The barrier for AI technology is fading away, bringing a lot of disruptions.However, Data is not cheapWhen technology becomes free, data is the King. Companies with high volume of data enjoys high valuation from investors because data can create a barrier that is almost impossible for its competitors to catch up.However not all the data can be efficiently used for machine learning models. Data scientists will probably share the same experience: most of the time for a machine learning project will be spent on cleansing the data rather than building the model. And if we are not so lucky that we don't have enough high-quality label data then we will need even more time to crawl and label the data manually.IN MY VIEW
< Page 7 | Page 9 >