{"id":2276,"date":"2023-05-01T14:28:50","date_gmt":"2023-05-01T14:28:50","guid":{"rendered":"https:\/\/leyton.majjane.agency\/ca\/post\/bidirectional-encoder-representations-from-transformers\/"},"modified":"2023-05-01T14:28:50","modified_gmt":"2023-05-01T14:28:50","slug":"bidirectional-encoder-representations-from-transformers","status":"publish","type":"article","link":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/","title":{"rendered":"Bidirectional Encoder Representations from Transformers"},"content":{"rendered":"<p>The technological objective of <strong>Bidirectional Encoder Representations from Transformers<\/strong> <strong>(BERT)<\/strong> was to explore the geometry of BERT\u2019s internal representations of linguistic information including syntactic features and semantic features. BERT model is a new technique for NLP(natural language processing) pre-training developed by <strong>Google AI Language team<\/strong> which uses a multi-layer bidirectional transformer encoder and two unsupervised pre-trained tasks including masked LM and NSP(next sentence prediction). Instead of the traditional left-to-right or right-to-left model, <strong><a href=\"https:\/\/leyton.com\/ca\/ict-companies\/\">BERT<\/a><a href=\"https:\/\/leyton.com\/ca\/software-en\/\"> <\/a><\/strong>pre-trains unlabeled text by jointly conditioning on both left and right context in all layers. As a result,<a href=\"https:\/\/leyton.com\/ca\/ict-companies\/\"> <strong>BERT <\/strong><\/a>is capable of a wide range of tasks, such as question answering and language inference.<\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>When reviewing the previous studies, related works on similar neural nets such as CNN and Word2Vec have led to some knowledge limitations about the current study such as which linguistic features are translated into geometry representations and some hypotheses like grammatical information may be represented via directions in space and attention matrices may encode important relations between words. <\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>Furthermore, the research focused on the geometric representation of the entire parse trees by Hewitt and Manning points to two knowledge limitations on the technology, one is the possibility of discovering other examples of intermediate representations, and another is how these internal representations decompose.<\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>There were a number of specific technological obstacles that drove the investigations described further. First, <strong><a href=\"https:\/\/leyton.com\/ca\/ict-companies\/\">BERT<\/a><a href=\"https:\/\/leyton.com\/ca\/software-en\/\"> <\/a><\/strong>is a newly released <strong>natural language processing<\/strong> (NLP) framework that Google calls it the biggest leap forward in five years. Unlike traditional neural nets such as CNN or RNN which have sufficient previous works to refer to, <strong><a href=\"https:\/\/leyton.com\/ca\/ict-companies\/\">BERT\u2019s<\/a><a href=\"https:\/\/leyton.com\/ca\/software-en\/\"> <\/a><\/strong>transformer architecture had been a largely underexplored domain with many untapped potentials. Therefore it is difficult to find some existing technique or methodology corresponding to this study.<\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<figure class=\"wp-block-image aligncenter size-full\"><img decoding=\"async\" src=\"https:\/\/leyton.com\/ca\/wp-content\/blogs.dir\/3\/files\/2023\/04\/Bert-2B-1-1.jpg\" alt=\"BERT\" class=\"wp-image-22200\" \/><\/figure>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>The second technological shortcoming in this study was, to finally visualize the internal geometry, they had to deal with high-dimensional space. Begin with finding theoretical explanations to prove the existence, then try some techniques to project down to 2-dimensions which provide understandable images. <\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>As a result, the experiments need to be carefully designed and the technique for projection requires cautiously selection. For instance, while visualizing tree embeddings for the geometry of syntax, the study uses PCA due to its easiness of interpreting. However, during the visualization of word sense, they use UMAP since PCA tends to lose some of the subtitles while UMAP has speed-ups and the ability to better preserve the data below the global structure.<\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>Studies of <strong>NLP <\/strong>to improve user\u2019s searching experience have been the subject of much research in recent years. This work discussed the syntactic representation in attention matrices and directions in space representing dependency relation. They also proposed a mathematical justification for the squared-distance tree embedding and visualized tree embedding to prove syntactic representation has a quantitative aspect. Further, they investigate how mistakes in word sense disambiguation may correspond to changes in internal geometry representation for word meaning. They also explored that the internal geometry of <strong><a href=\"https:\/\/leyton.com\/ca\/ict-companies\/\">BERT<\/a> <\/strong>may be split into multiple linear subspaces to fit different representations together.<\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>The result presented in this paper will bring significant influences on <strong>natural language processing tasks<\/strong>. One of the most exciting findings in this study is the ability of internal geometry to break into separate linear subspaces for different syntactic and semantic information. This kind of decomposition implies there may have other meaningful subspaces and represent other types of linguistic features. <\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>Another potential avenue of exploration is when discovering the linear transformation for embedding subspace, instead of the final layer, the result suggests there is more semantic information in the geometry of earlier-layer embedding, which can achieve higher state-of-the-art accuracy. Moreover, the result of the concatenation experiment points to a potential failure mode of attention-based models. These valuable results on internal geometry inspire people to achieve a deeper understanding of <strong>transformer <\/strong>architecture and shed light on the improvement of <strong><a href=\"https:\/\/leyton.com\/ca\/ict-companies\/\">BERT\u2019s<\/a><a href=\"https:\/\/leyton.com\/ca\/software-en\/\"> <\/a><\/strong>architecture.<\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>Companies that are innovating <a href=\"https:\/\/leyton.com\/ca\/ict-companies\/\">software<\/a><a href=\"https:\/\/leyton.com\/ca\/software-en\/\" target=\"_blank\" rel=\"noreferrer noopener\"> <\/a>technologies that can elevate our daily lives are likely to be eligible for several funding programs including <a href=\"https:\/\/leyton.com\/ca\/government-grants\/\" target=\"_blank\" rel=\"noreferrer noopener\">government grants,<\/a> and <a href=\"https:\/\/leyton.com\/ca\/tax-credits\/sred-already-claiming\/\" target=\"_blank\" rel=\"noreferrer noopener\">SR&amp;ED<\/a>.<\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<p>Want to learn about funding opportunities for your project? <a href=\"https:\/\/leyton.com\/ca\/contact-us\/\" target=\"_blank\" rel=\"noreferrer noopener\"><strong>Schedule a free consultation with one of our experts today!<\/strong><\/a><\/p>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n<div class=\"wp-block-buttons\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/leyton.com\/ca\/contact-us\/\" target=\"_blank\" rel=\"noreferrer noopener\">Contact Us<\/a><\/div>\n<\/div>\n<div style=\"height:30px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n","protected":false},"excerpt":{"rendered":"<p>The technological objective of Bidirectional Encoder Representations from Transformers (BERT) was to explore the geometry of BERT\u2019s internal representations of linguistic information including syntactic features and semantic features. BERT model is a new technique for NLP(natural language processing) pre-training developed by Google AI Language team which uses a multi-layer bidirectional transformer encoder and two unsupervised [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2968,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[230,110],"tags":[232,114,590,115],"expertise":[774],"class_list":["post-2276","article","type-article","status-publish","format-standard","has-post-thumbnail","hentry","category-software","category-sred","tag-funding","tag-innovation-en","tag-software","tag-sred-en","expertise-sred-tax-credits"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v28.1 (Yoast SEO v28.1) - https:\/\/yoast.com\/product\/yoast-seo-premium-wordpress\/ -->\n<title>Bidirectional Encoder Representations from Transformers - Leyton Canada<\/title>\n<meta name=\"description\" content=\"What is the objective of a Bidirectional Encoder Representations from Transformers? The BERT model is a new technique developed by Google\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Bidirectional Encoder Representations from Transformers\" \/>\n<meta property=\"og:description\" content=\"What is the objective of a Bidirectional Encoder Representations from Transformers? The BERT model is a new technique developed by Google\" \/>\n<meta property=\"og:url\" content=\"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/\" \/>\n<meta property=\"og:site_name\" content=\"Leyton Canada\" \/>\n<meta property=\"og:image\" content=\"https:\/\/leyton.com\/wp-content\/blogs.dir\/3\/files\/2023\/04\/bert.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1920\" \/>\n\t<meta property=\"og:image:height\" content=\"655\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"4 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/\",\"url\":\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/\",\"name\":\"Bidirectional Encoder Representations from Transformers - Leyton Canada\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/leyton.com\\\/wp-content\\\/blogs.dir\\\/3\\\/files\\\/2023\\\/04\\\/bert.jpg\",\"datePublished\":\"2023-05-01T14:28:50+00:00\",\"description\":\"What is the objective of a Bidirectional Encoder Representations from Transformers? The BERT model is a new technique developed by Google\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/#primaryimage\",\"url\":\"https:\\\/\\\/leyton.com\\\/wp-content\\\/blogs.dir\\\/3\\\/files\\\/2023\\\/04\\\/bert.jpg\",\"contentUrl\":\"https:\\\/\\\/leyton.com\\\/wp-content\\\/blogs.dir\\\/3\\\/files\\\/2023\\\/04\\\/bert.jpg\",\"width\":1920,\"height\":655},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/en\\\/insights\\\/articles\\\/bidirectional-encoder-representations-from-transformers\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/leyton.com\\\/ca\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Bidirectional Encoder Representations from Transformers\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/#website\",\"url\":\"https:\\\/\\\/leyton.com\\\/ca\\\/\",\"name\":\"Leyton\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/leyton.com\\\/ca\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/#organization\",\"name\":\"Leyton\",\"url\":\"https:\\\/\\\/leyton.com\\\/ca\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/leyton.com\\\/wp-content\\\/blogs.dir\\\/3\\\/files\\\/2026\\\/01\\\/logo.svg\",\"contentUrl\":\"https:\\\/\\\/leyton.com\\\/wp-content\\\/blogs.dir\\\/3\\\/files\\\/2026\\\/01\\\/logo.svg\",\"width\":108,\"height\":44,\"caption\":\"Leyton\"},\"image\":{\"@id\":\"https:\\\/\\\/leyton.com\\\/ca\\\/#\\\/schema\\\/logo\\\/image\\\/\"}}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"Bidirectional Encoder Representations from Transformers - Leyton Canada","description":"What is the objective of a Bidirectional Encoder Representations from Transformers? The BERT model is a new technique developed by Google","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/","og_locale":"en_US","og_type":"article","og_title":"Bidirectional Encoder Representations from Transformers","og_description":"What is the objective of a Bidirectional Encoder Representations from Transformers? The BERT model is a new technique developed by Google","og_url":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/","og_site_name":"Leyton Canada","og_image":[{"width":1920,"height":655,"url":"https:\/\/leyton.com\/wp-content\/blogs.dir\/3\/files\/2023\/04\/bert.jpg","type":"image\/jpeg"}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"4 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/","url":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/","name":"Bidirectional Encoder Representations from Transformers - Leyton Canada","isPartOf":{"@id":"https:\/\/leyton.com\/ca\/#website"},"primaryImageOfPage":{"@id":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/#primaryimage"},"image":{"@id":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/#primaryimage"},"thumbnailUrl":"https:\/\/leyton.com\/wp-content\/blogs.dir\/3\/files\/2023\/04\/bert.jpg","datePublished":"2023-05-01T14:28:50+00:00","description":"What is the objective of a Bidirectional Encoder Representations from Transformers? The BERT model is a new technique developed by Google","breadcrumb":{"@id":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/#primaryimage","url":"https:\/\/leyton.com\/wp-content\/blogs.dir\/3\/files\/2023\/04\/bert.jpg","contentUrl":"https:\/\/leyton.com\/wp-content\/blogs.dir\/3\/files\/2023\/04\/bert.jpg","width":1920,"height":655},{"@type":"BreadcrumbList","@id":"https:\/\/leyton.com\/ca\/en\/insights\/articles\/bidirectional-encoder-representations-from-transformers\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/leyton.com\/ca\/"},{"@type":"ListItem","position":2,"name":"Bidirectional Encoder Representations from Transformers"}]},{"@type":"WebSite","@id":"https:\/\/leyton.com\/ca\/#website","url":"https:\/\/leyton.com\/ca\/","name":"Leyton","description":"","publisher":{"@id":"https:\/\/leyton.com\/ca\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/leyton.com\/ca\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/leyton.com\/ca\/#organization","name":"Leyton","url":"https:\/\/leyton.com\/ca\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/leyton.com\/ca\/#\/schema\/logo\/image\/","url":"https:\/\/leyton.com\/wp-content\/blogs.dir\/3\/files\/2026\/01\/logo.svg","contentUrl":"https:\/\/leyton.com\/wp-content\/blogs.dir\/3\/files\/2026\/01\/logo.svg","width":108,"height":44,"caption":"Leyton"},"image":{"@id":"https:\/\/leyton.com\/ca\/#\/schema\/logo\/image\/"}}]}},"_links":{"self":[{"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/article\/2276","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/article"}],"about":[{"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/types\/article"}],"author":[{"embeddable":true,"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/comments?post=2276"}],"version-history":[{"count":0,"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/article\/2276\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/media\/2968"}],"wp:attachment":[{"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/media?parent=2276"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/categories?post=2276"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/tags?post=2276"},{"taxonomy":"expertise","embeddable":true,"href":"https:\/\/leyton.com\/ca\/wp-json\/wp\/v2\/expertise?post=2276"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}