ML Data Linguist , AWS AI Data

DESCRIPTION:

Amazon Web Services (AWS) is looking for a data associate to help with annotations and data analysis. As part of the AiData Team at AWS you will responsible for delivering high-quality training data to ensure the best performance of the AWS machine learning systems. Our goal is to produce the highest quality training data in the industry and to delight our customers by improving human language understanding and natural language processing.

The Bedrock team is a team of data linguists who primarily support the training of different models in the AWS generative AI platform. We are specialized in text-based data annotation, writing for ML model training, and toxic content evaluation. Some of the aspects of ML development that the Bedrock team works with include Responsible AI, Reinforcement Learning from Human Feedback, Supervised Fine Tuning, and Human Content Evaluation. Our team represents a great array of experience in the field of linguistics, including sociolinguistics, computational linguistics, conversation analysis, syntax-semantics, linguistic typology, ESL and foreign languages, as well as translation.

AWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon's Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS's services and features apart in the industry. As a member of the UC organization, you'll support the development and management of Compute, Database, Storage, Internet of Things (Iot), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services.

Key job responsibilities
* Build a thorough understanding of data collection and annotation guidelines and various annotation tools.
* Annotate text data, identifying linguistic categories based on detailed annotation and adhering to guidelines.
* Perform annotation related tasks; you participate in data generation, collection and quality assurance tasks
* Collaborate with other ML Data Linguists to resolve data ambiguities and annotation disagreements.
* Dive deep into the data to perform qualitative error trend analysis.
* Provide feedback to Language Engineers and Scientists on tool improvements and annotation processes.
* Diving deep into issues and implement solutions independently
* Contribute to process improvements to reduce handling time and improve resource output.
* Develop a variety of language artifacts crucial for model development such as datasets for training and evaluation.

We are open to hiring candidates to work out of one of the following locations:

Virtual Location - USA

BASIC QUALIFICATIONS:

* Bachelor's degree in a relevant field, such as Linguistics, Communications, a foreign language, or other language or data-related disciplines.
* 6 months of experience with natural language data labeling, data annotation, linguistic annotation or other forms of data markup, and/or teaching experience.
* Proficient in Spanish, French, German, Portuguese, Japanese, Korean, or another foreign language.
* Experience identifying linguistic ambiguity and annotation inaccuracies in data.
* Ability to strictly adhere to annotation guidelines and identify basic parts of speech.
* Strong organizational skills and detail-oriented
* Ability to communicate well and actively listen with other data associates on a team.
* Ability to deliver high quality results under tight deadlines.
* Comfortable working in a fast paced, collaborative work environment.
* Willingness to support several projects at one time, and to accept re-prioritization as necessary.

PREFERRED QUALIFICATIONS:

* 1+ years of experience in the language data annotation.
* Ability to quickly learn new data annotation guidelines, technical concepts, and softwares.
* Depth and breadth of knowledge in linguistic theory and/or applied linguistics.
* Familiarity with common text processing tools.
* Passion for language, linguistics, human language technology and AI.
* Familiarity with json, yaml, xml or other forms of text markup.
* Ability to work in different operating systems (Windows, MacOS, or Linux).
* Ability to navigate a Unix terminal and use common command line tools
* Knowledge of Python, Java or any other scripting language is a plus.

Amazon is committed to a diverse and inclusive workplace. Amazon is an equal opportunity employer and does not discriminate on the basis of race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status. For individuals with disabilities who would like to request an accommodation, please visit https://www.amazon.jobs/en/disability/us.

Pursuant to the Los Angeles Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our compensation reflects the cost of labor across several US geographic markets. The base pay for this position ranges from $32,700/year in our lowest geographic market up to $70,000/year in our highest geographic market. Pay is based on a number of factors including market location and may vary depending on job-related knowledge, skills, and experience. Amazon is a total compensation company. Dependent on the position offered, equity, sign-on payments, and other forms of compensation may be provided as part of a total compensation package, in addition to a full range of medical, financial, and/or other benefits. For more information, please visit https://www.aboutamazon.com/workplace/employee-benefits. This position will remain posted until filled. Applicants should apply via our internal or external career site.

Apply Now

	Organisation	Amazon
	Job Area	Software
	Industry	Internet
	Location	Remote
	Country	United States
	Salary	Not available